{
  "id": 8058129,
  "title": "Android Bench 2.0 focuses on long-horizon tasks, agent evaluations",
  "url": "https://urgent.news/2026/09/17/android-bench-2-0-focuses-on-long-horizon-tasks-agent-evaluations",
  "topic": "ai",
  "section": "AI",
  "published": "2026-09-17T16:00:00.000Z",
  "source": {
    "name": "9to5Google",
    "slug": "9to5google",
    "url": "https://9to5google.com/2026/09/17/android-bench-2-0/"
  },
  "original_language": "en",
  "account": "Google has unveiled Android Bench 2.0, a development tool aimed at assessing complex AI tasks beyond simple bug fixes or minor updates. This new version zeroes in on long-horizon tasks that typically consume multiple days or even a week for human engineers to accomplish, such as creating new features, constructing apps from scratch, and converting cross-platform software to Android.\n\nTo better evaluate the performance of AI agents handling these intricate assignments, Android Bench 2.0 employs a more sophisticated scoring system. It now provides continuous scores instead of simple pass or fail ratings, taking into account factors like functionality, visual appearance, and the avoidance of regressions. The system also imposes penalties for any deviation from the evaluation guidelines or structural constraints.\n\nTo gauge the prowess of various AI models, Google has benchmarked them using Android Bench 2.0. The results show that GPT-6 Astra leads the pack with a 28% pass rate, while other models like Gemini 3.7/3.8 Flash, OpenAI's GPT-6, Anthropic's Fable 5.1, Kimi K3, and Qwen 3.8 Max trail behind, with scores generally in the high 90s from the previous benchmark's perspective.\n\nThe report highlights that converting cross-platform apps to Android remains a demanding challenge, as no AI model has achieved a flawless 100% pass rate. In fact, even the most advanced frontier models typically manage to complete around 80% of such tasks. During testing, Android Bench 2.0 ran agents provided by the respective model creators, such as Gemini 3.8 Flash on Google's Antigravity platform and GPT-5.6 Sol with Codex.\n\nGoogle's team believes that the design of these agent evaluations significantly influences the overall developer experience. As such, they plan to incorporate a wider range of models and agent combinations in future iterations of Android Bench 2.0. Stay tuned for more updates on this development at 9to5Google on YouTube.",
  "summary": "Google’s development of Android Bench continues today with a version 2.0 that reflects how AI can handle more complex development tasks. more…",
  "key_points": [],
  "editors_take": null,
  "illustration": null,
  "coverage": {
    "outlets": 1,
    "also_reported_by": []
  },
  "ai_generated": true,
  "disclaimer": "Summaries, key points and the editor’s take are written by software from other outlets’ reporting and may contain errors — always check the linked original."
}