Google Releases Android Bench 2.0 With Long-Horizon Tasks and Agentic Evaluation
Tech

Google Releases Android Bench 2.0 With Long-Horizon Tasks and Agentic Evaluation

TechNews Editorial
TechNews EditorialOct 10, 2026 · 2 min read
Share

Why it matters

The update matters because it provides developers with a structured benchmark to measure how AI agents handle complex, multi-day Android development tasks.

The facts

  • Google released Android Bench 2.0 with long-horizon tasks, agentic evaluation, and continuous scoring.
  • The update shifts from binary pass or fail scoring to a nuanced completion rate evaluating functionality and visual fidelity.
  • Claude Opus 5.5 topped the leaderboard with a 32% long-horizon task pass rate followed by GPT 6 Astra at 28%.

Google has released Android Bench 2.0. This is a major update to its benchmark framework for evaluating AI models and agents on Android development tasks. The update introduces long-horizon tasks, agent-based evaluation, and continuous scoring to better assess performance on complex, multi-step development tasks.

Launched a few months ago, Android Bench evaluates AI models against a set of common development tasks. It incorporates Android best practices in areas such as permissions, navigation, and connectivity. With the latest release, Google has expanded the benchmark to cover a broader range of development tasks.

Google added complex long-horizon tasks

Google stated that they are releasing the first set of long-horizon tasks, which are tasks of great complexity that take an engineer multiple days or even a week to complete. They are also introducing agentic evaluation, starting with agents from corresponding model providers.

While the original version focused on incremental changes to existing repositories, Android Bench 2.0 introduces long-horizon tasks that Google describes as work an engineer might take multiple days or even a week to complete. These tasks include upgrading dependencies, adding new features, building apps from scratch, or converting a cross-platform app to Android.

An automated evaluator rejects an entire test report for one failure, while a revised evaluator preserves credit for the many successful requirements.
Illustration: AI & Tech News

Binary scoring changed to completion rates

One major change in version 2.0 is the shift from binary pass or fail evaluation to a more nuanced scoring system. Previously, a complex task could be marked as failed because of a single failing edge-case assertion, even if the agent had successfully met dozens of other requirements.

Google stated they calculate this completion rate through a combination of factors like functionality, visual fidelity, and avoiding regressions. They also apply objective scoring penalties for deviations from evaluation instructions or structural constraints.

Read nextGoogle Releases AI Edge Foresight Local-First Meeting Note-Taker

Models struggle with runtime validation

The results provided by Android Bench 2.0 also help identify which tasks are more likely to succeed with AI assistance. For example, Google reports that AI does a better job at writing new code rather than refactoring existing code, which can reflect the fact that refactoring and migrations require an understanding of the architectural complexity of the codebase.

Similarly, AI performs well on several well-established, deterministic transformations, even in larger codebases. Examples include converting Java to Kotlin, swapping Retrofit for Ktor, or introducing a ViewModel layer.

On the contrary, in a number of cases models still struggle, including with tasks requiring runtime validation, involving breaking framework changes, or running into knowledge gaps with unreleased libraries. In particular, the best-in-class model achieves only an 80% completion rate in porting cross-platform apps to Android, which remains overall an open challenge.

The updated Android Bench 2.0 dashboard includes recent models, such as Gemini 3.8 Flash, Gemini 3.7 Flash, OpenAI GPT-6, Anthropic Fable 5.1, Kimi K3, and Qwen 3.8 Max. At the time of the article, Claude Opus 5.5 was at the top of the leaderboard with a 32% long-horizon task pass rate, followed by GPT 6 Astra at 28%.

Newsletter

Get the best AI & tech news daily

A concise daily digest. Unsubscribe anytime.

We use your email only to send this newsletter.

Keep reading