PulseAugur
EN
LIVE 18:18:33

Google's Android Bench adds new LLMs; Fable 5 leads, Gemini lags

Google has updated its Android Bench benchmark for evaluating large language models (LLMs) in Android development tasks. The updated leaderboard includes eight new models, such as Claude Fable 5, Claude Sonnet 5, and Qwen 3.7 Max. Notably, Claude Fable 5 leads in accuracy at 84.5 percent, while Google's own Gemini 3.1 Pro ranks fifth. The benchmark also highlights significant cost differences between models, with Fable 5 and GPT 5.5 being the most expensive to run. AI

IMPACT Provides developers with updated performance data to select the best LLMs for Android coding tasks.

RANK_REASON Update to an existing benchmark tool for LLMs in software development.

Read on Ars Technica — AI →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

Google's Android Bench adds new LLMs; Fable 5 leads, Gemini lags

How we ranked this

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Tool
Update to an existing benchmark tool for LLMs in software development.
Source corroboration
Single-source cluster
Only one publisher covered this so far. Single-source stories can still rank when the publisher is high-authority, but they lack cross-source corroboration.
Topics
product, model release
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
80 days old
Aged out of breaking-news scoring windows; ranking reflects the durable signal from the full source set.

Full methodology in our editorial standards.

COVERAGE [1]

  1. Ars Technica — AI TIER_1 English(EN) · Ryan Whitwam ·

    Google updates Android Bench with new LLMs, but Gemini still lags behind

    Android Bench is evolving, and developers can help guide that process.