Android Bench 2 Adds Support for Long-Horizon Tasks, Agentic Evaluation, and Continuous Scoring
Google has released Android Bench 2.0, a major update to its benchmark framework for evaluating AI models and agents on Android development tasks. The update introduces long-horizon tasks (LHTs), agent-based evaluation, and continuous scoring to better assess performance on complex, multi-step development tasks. By Sergio De Simone
Google has launched Android Bench 2.0, a significant upgrade to its benchmark framework designed for assessing AI models and agents in Android development. This update brings long-horizon tasks (LHTs), agentic evaluation, and continuous scoring to the framework, aiming to provide a more comprehensive evaluation of AI performance on complex, multi-step development challenges.
Launched a few months ago, Android Bench initially assessed AI models on a series of common development tasks, adhering to Android best practices in areas such as permissions, navigation, and connectivity.
The latest version now includes the first set of LHTs, which are tasks of considerable complexity that typically require an engineer to spend multiple days or even a week to complete. In addition, Android Bench 2.0 introduces agentic evaluation, with initial evaluations featuring agents from the corresponding model providers. While the previous version concentrated on incremental modifications to existing codebases, Android Bench 2.0 introduces LHTs that encompass significant tasks an engineer might take days or even a week to accomplish.
These tasks include updating dependencies, introducing new features, constructing apps from the ground up, or transforming a cross-platform app to Android.
A key change in version 2.0 is the transition from a binary pass/fail evaluation to a more sophisticated scoring system. The previous system would flag a complex task as failed due to a single failing edge-case assertion, even if the AI had succeeded in meeting numerous other requirements. Android Bench 2.0 calculates completion rates through a combination of factors including functionality, visual fidelity, and the avoidance of regressions.
Moreover, objective scoring penalties are applied for any deviations from the evaluation instructions or structural constraints.
The results from Android Bench 2.0 also shed light on which tasks are more likely to be successful with AI assistance. Google notes that AI excels at generating new code rather than refactoring existing code, a distinction reflecting the fact that refactoring and migrations necessitate understanding the architectural complexity of the codebase.
Similarly, AI performs well on several well-established, deterministic transformations, such as converting Java to Kotlin, swapping Retrofit for Ktor, or incorporating a ViewModel layer. However, in some cases, AI models still struggle, particularly with tasks requiring runtime validation (such as missing dependency injection graphs), involving breaking framework changes, or dealing with knowledge gaps regarding unreleased libraries.
Specifically, the best-in-class model achieves an 80% completion rate in porting a cross-platform app to Android, which remains an open challenge. The updated Android Bench 2.0 dashboard now features recent models like Gemini 3.8 Flash, Gemini 3.7 Flash, OpenAI GPT-6, Anthropic Fable 5.1, Kimi K3, and Qwen 3.8 Max. As of the article's publication, Claude Opus 5.5 led the leaderboard with a 32% LHT pass rate, followed by GPT 6 Astra at 28%.
Written by urgent.news from InfoQ's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.