Urgent.News

What's breaking now, across thousands of outlets.

Culture

Echoverse: Deep, evolving environments for computer-use agents

Computer-use AI agents struggle with multi-step workflows like email and customer support. Echoverse trains agents in realistic environments rather than simply providing more training tasks, helping them improve as the tasks, tests, and environments evolve. The post Echoverse: Deep, evolving environments for computer-use agents appeared first on Microsoft Research .

Echoverse: Deep, evolving environments for computer-use agents

At the core of the Echoverse project are deep, evolving environments designed for training computer-use agents. By focusing on fidelity over sheer count, the team built twelve training worlds, ten deep domain worlds and two capability worlds, each focused on a specific control in various forms. These worlds accurately replicate real application behavior, come with realistic data, and maintain consistent state across screens and users.

Trained on all twelve worlds, a 9B model achieved a notable improvement in its base score, from 36.5% to 67.1%, coming close to GPT-5.4's performance. This experiment highlighted the importance of high simulation fidelity and demonstrated how shallow worlds can hinder an agent's development.

The model often struggled with the same challenging UI elements, such as date pickers and nested filters, which were targeted in the Echoverse training. By drilling these controls in various forms, the model learned to operate them effectively in new domains. Furthermore, co-evolving the model, the world, and the verifier led to improvements in all three components, with the model climbing as the world grew more accurate and tasks more complex.

Reinforcement learning against the worlds pushed the agent past imitation, using the grounded verifier as a reward to teach the agent to reach goals in fewer steps. Four of the worlds, along with their code, data, and grounded graders, have been released on GitHub and Hugging Face to support research on high-fidelity computer-use worlds.

This approach to training computer-use agents differs from simply adding more environments, as the real leverage comes from a loop that continuously improves the existing worlds. By treating the building of the environment and training of the model as one process, rather than two separate stages, the loop compiles and yields better results than ordinary fine-tuning.

Written by urgent.news from Microsoft Research's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at microsoft.com →

More in Culture

COR raises $30M led by FTV Capital

Argentine startup COR raised $30M in a funding round led by FTV Capital. As part of the investment,… The post COR raises $30M led by FTV Capital appeared first on LatamList .

More from Thursday 30 July →