{
  "id": 3767603,
  "title": "“Posterity will find it ludicrous”: Sai agent hits 73% on OSWorld 2.0 performing routine (but necessary) work",
  "url": "https://urgent.news/2026/08/27/posterity-will-find-it-ludicrous-sai-agent-hits-73-on-osworld-2-0",
  "topic": "ai",
  "section": "AI",
  "published": "2026-08-27T15:34:11.000Z",
  "source": {
    "name": "The New Stack",
    "slug": "the-new-stack",
    "url": "https://thenewstack.io/sai-agent-osworld-benchmark/"
  },
  "original_language": "en",
  "account": "Sai, a computer agent designed by Simular for everyday, lengthy professional tasks, has achieved a 73% success rate on the OSWorld 2.0 benchmark. This performance surpasses GPT-5.6 Sol at 62.57% and Opus 5 at 70.57%, indicating that Sai is currently the most efficient model in its class. Unlike other models, Sai focuses on real-world workplace tasks and functions, such as recruitment outreach, invoicing validation, and researching news, rather than achieving industry accolades through unfettered throughput. Simular's co-founder and CTO, Jiachen Yang, believes that computer agents for routine work should not be prohibitively expensive, stating that \"posterity will find it ludicrous that people are still building models that way right now.\" The underlying technology behind Sai's efficiency is its neurosymbolic method, which combines the flexible exploratory abilities of neural networks with the precision of symbolic code. This allows Sai to encode solved tasks into reusable code, making the processing of the hundredth invoice cheaper than the first one. Sai's efficiency is further improved by caching and memory efficiency techniques, including adaptive summarization and maintaining a constant prompt prefix for as long as possible. Additionally, Sai uses fewer model calls on average than pure models by taking more actions per turn and employing Simulang code as a symbolic language for planning and executing longer subtasks.",
  "summary": "Sai, a computer agent built by Simular, has achieved a 73% success rate on OSWorld 2.0, in a benchmark update The post “Posterity will find it ludicrous”: Sai agent hits 73% on OSWorld 2.0 performing routine (but necessary) work appeared first on The New Stack .",
  "key_points": [],
  "editors_take": null,
  "illustration": null,
  "coverage": {
    "outlets": 1,
    "also_reported_by": []
  },
  "ai_generated": true,
  "disclaimer": "Summaries, key points and the editor’s take are written by software from other outlets’ reporting and may contain errors — always check the linked original."
}