PulseAugur
EN
LIVE 00:40:30

AI Labs Not Gaming 'Pelican on Bicycle' Benchmark, Study Finds

An experiment was conducted to determine if AI labs are optimizing their models for a specific informal benchmark: generating an SVG of a pelican riding a bicycle. The study tested seven frontier models, including GPT-5.6 Terra and Claude Sonnet 5, by generating 1,008 SVGs across various animal-vehicle combinations. The results, analyzed using an LLM judge and Gemini 3.1 Flash-Lite, suggest that labs are not specifically 'pelicanmaxxing' their models, as performance on the pelican-bicycle prompt did not significantly outperform other combinations for the tested models. AI

IMPACT Suggests that current frontier models are not specifically optimized for niche, informal benchmarks, indicating a focus on broader capabilities.

RANK_REASON The cluster discusses an experiment and analysis of AI model performance on a specific, informal benchmark, which falls under research.

Read on Simon Willison →

AI-generated summary · Google Gemini · from 5 sources. How we write summaries →

AI Labs Not Gaming 'Pelican on Bicycle' Benchmark, Study Finds

How we ranked this

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Research
The cluster discusses an experiment and analysis of AI model performance on a specific, informal benchmark, which falls under research.
Source corroboration
5 independent sources
Strong cross-source corroboration — multiple independent publishers covered this within the clustering window.
Topics
model release, product
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
68 days old
Aged out of breaking-news scoring windows; ranking reflects the durable signal from the full source set.
Coverage growth since scoring
+1 source(s) since last score
New sources have picked up this story since our last re-score. Score will update on the next scoring pass.

Full methodology in our editorial standards.

COVERAGE [5]

  1. Simon Willison TIER_1 (CA) ·

    Are AI labs pelicanmaxxing?

    <p><strong><a href="https://dylancastillo.co/posts/pelicanmaxxing.html">Are AI labs pelicanmaxxing?</a></strong></p> Excellent piece of work by Dylan Castillo, who took a deep-dive into the frequently pondered question of whether the AI labs have been deliberately training models…

  2. Hacker News — AI stories ≥50 points TIER_1 (CA) · dcastm ·

    Are AI Labs Pelicanmaxxing?

  3. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    Are AI Labs Pelicanmaxxing? https:// dylancastillo.co/posts/pelican maxxing.html Comments: https:// news.ycombinator.com/item?id=4 9010129 # HackerNews # AI # L

    Are AI Labs Pelicanmaxxing? https:// dylancastillo.co/posts/pelican maxxing.html Comments: https:// news.ycombinator.com/item?id=4 9010129 # HackerNews # AI # Labs # Pelicanmaxxing # Tech # Trends # Innovation # MachineLearning

  4. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    Are AI Labs Pelicanmaxxing? https:// dylancastillo.co/posts/pelican maxxing.html # ai

    Are AI Labs Pelicanmaxxing? https:// dylancastillo.co/posts/pelican maxxing.html # ai

  5. Mastodon — mastodon.social TIER_1 English(EN) · [email protected] ·

    Are AI labs pelicanmaxxing? https://simonwillison.net/2026/Jul/22/are-ai-labs-pelicanmaxxing/#atom-everything # AI # Tech # OpenSource

    Are AI labs pelicanmaxxing? https://simonwillison.net/2026/Jul/22/are-ai-labs-pelicanmaxxing/#atom-everything # AI # Tech # OpenSource