Urgent.News

What's breaking now, across thousands of outlets.

Tech

Driving a real desktop without a VM, a sandbox, or stealing your cursor

Driving a real desktop without a VM, a sandbox, or stealing your cursor Every "computer use" agent demo works in a browser tab. The hard part is doing it on a real Windows machine — behind the user's live session — without hijacking the cursor or breaking the thing you're automating. Most agent frameworks stop at the browser. They drive a headless page or a containerised desktop and call it…

A desktop agent that can operate a real Windows machine without taking over the user's cursor or breaking the task at hand is OpenAmer. This open-source desktop agent (licensed under Apache-2.0) focuses on running behind the user's live session, without the need for a VM, container, or stealing the cursor. The key design decisions that enable this are:

1. Control the desktop, not take it over: The agent captures the target window's state and delivers synthetic input to it without competing for the physical pointer. This ensures the user's session maintains its own cursor and focus, allowing them to continue using the machine. The agent deals with coordinate space, focus, and timing challenges to ensure accurate input delivery.

2. No VM, no container, no cloud: Unlike screenshot-driven agents that run in a disposable VM, OpenAmer runs as a process on the host, using the user's own credentials and session. This allows it to interact with the user's real files, logins, and applications, while also treating security trade-offs explicitly - capabilities are bounded, actions crossing boundaries are surfaced, and outcomes are logged for verification.

3. Drive a real browser profile, not a throwaway one: Web work requires the user's logins, so the agent controls a persistent Chrome profile using the DevTools protocol (CDP). The agent respects the user's authenticated identity and verifies results against the actual world, rather than relying on the model's self-report.

4. Verify against the world, not the model's self-report: The agent refuses to trust the model's answer to "did that succeed?" and instead checks the outcome ledger, plus re-reads the actual world when necessary. This ensures the agent's actions are verified and trustworthy, allowing it to run unattended.

Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at dev.to →

More in Tech

Wildhunt_AI

This is a submission for the Hacktoberfest Open-Source AI Challenge Week 1: Touch Grass What I Built WildHunt AI is a privacy-first, open-source AI scavenger hunt that turns your phone into a guide…

  • WildHunt AI encourages phone users to explore the physical world.
  • App functions as a Progressive Web App, accessible via web browser.
  • AI verifies photos against specific missions, enhancing exploration.

Why JavaScript Says 5 + 5 = 55 (and How to Fix It)

You build a small calculator in JavaScript. Two input boxes, one button. You type 5 in the first box, 5 in the second, click the button, and the page says 55 . Not 10.

  • JavaScript treats input values as text (strings) in calculations
  • Use Number() for accurate numeric calculations, including decimals
  • Check for invalid inputs using isNaN() and trim whitespace

Sobrevivendo ao Inevitável: Engenharia do Caos na Prática com o ChaosEngineeringMasterDeck

Se você trabalha com arquiteturas modernas em nuvem e microsserviços, provavelmente já se deparou com a clássica citação de Werner Vogels, CTO da Amazon: "Failures are a given and everything will…

  • Modern cloud architectures face increasing complexity with interconnected microservices.
  • Chaos Engineering tests systems to build resilience against unexpected production issues.

More from Sunday 11 October →