Driving a real desktop without a VM, a sandbox, or stealing your cursor
Driving a real desktop without a VM, a sandbox, or stealing your cursor Every "computer use" agent demo works in a browser tab. The hard part is doing it on a real Windows machine — behind the user's live session — without hijacking the cursor or breaking the thing you're automating. Most agent frameworks stop at the browser. They drive a headless page or a containerised desktop and call it…
A desktop agent that can operate a real Windows machine without taking over the user's cursor or breaking the task at hand is OpenAmer. This open-source desktop agent (licensed under Apache-2.0) focuses on running behind the user's live session, without the need for a VM, container, or stealing the cursor. The key design decisions that enable this are:
1. Control the desktop, not take it over: The agent captures the target window's state and delivers synthetic input to it without competing for the physical pointer. This ensures the user's session maintains its own cursor and focus, allowing them to continue using the machine. The agent deals with coordinate space, focus, and timing challenges to ensure accurate input delivery.
2. No VM, no container, no cloud: Unlike screenshot-driven agents that run in a disposable VM, OpenAmer runs as a process on the host, using the user's own credentials and session. This allows it to interact with the user's real files, logins, and applications, while also treating security trade-offs explicitly - capabilities are bounded, actions crossing boundaries are surfaced, and outcomes are logged for verification.
3. Drive a real browser profile, not a throwaway one: Web work requires the user's logins, so the agent controls a persistent Chrome profile using the DevTools protocol (CDP). The agent respects the user's authenticated identity and verifies results against the actual world, rather than relying on the model's self-report.
4. Verify against the world, not the model's self-report: The agent refuses to trust the model's answer to "did that succeed?" and instead checks the outcome ledger, plus re-reads the actual world when necessary. This ensures the agent's actions are verified and trustworthy, allowing it to run unattended.
Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.