{"id":1179908,"date":"2026-07-30T10:00:00","date_gmt":"2026-07-30T17:00:00","guid":{"rendered":"https:\/\/www.microsoft.com\/en-us\/research\/?p=1179908"},"modified":"2026-07-30T10:58:23","modified_gmt":"2026-07-30T17:58:23","slug":"echoverse-deep-evolving-environments-for-computer-use-agents","status":"publish","type":"post","link":"https:\/\/www.microsoft.com\/en-us\/research\/blog\/echoverse-deep-evolving-environments-for-computer-use-agents\/","title":{"rendered":"Echoverse: Deep, evolving environments for computer-use agents"},"content":{"rendered":"\n

Scaling fidelity over sheer count, targeting the capabilities agents actually lack, and evolving with the models they train.<\/h2>\n\n\n\n
\"Diagram<\/figure>\n\n\n\n
\n\t\n\t
\n\t\t
\n\t\t\t
\n
\n

At a glance<\/h2>\n\n\n\n

We built twelve training worlds for computer-use agents: ten deep domain worlds and two capability worlds, each drilling a single control rendered in many forms (date pickers and nested filters). Depth is what makes them worth training on: these worlds reproduce an application\u2019s real behavior, come seeded with realistic data, and keep state coherent across screens and users. Trained on all twelve, a 9B model nearly doubles its base score (36.5% to 67.1%), coming within fourteen points of GPT-5.4. The experiment taught us several lessons: <\/p>\n\n\n\n