PulseAugur
EN
LIVE 09:28:09

LLM benchmark: Pelicans on bikes show rapid progress over two years

Over the past two years, Simon Willison has been using a unique benchmark to track the progress of large language models: generating an SVG of a pelican riding a bicycle. Initially, models struggled with the task, producing abstract shapes and nonsensical interpretations. However, by mid-2025, models like DeepSeek R1 and Nvidia's offerings showed significant improvement, with some even beginning to "editorialise" by refusing the prompt or adding creative elements. A tournament using GPT-4.1 mini as a judge revealed Gemini 2.5 Pro preview as a top performer, highlighting the rapid advancements in LLM capabilities for code generation and image synthesis. AI

IMPACT Demonstrates rapid LLM progress in code generation and nuanced task interpretation, potentially influencing future benchmark design.

RANK_REASON The item describes a novel benchmark for evaluating LLMs and tracks their progress over time, including specific model performance and evolution. [lever_c_demoted from research: ic=1 ai=1.0]

Read on dev.to — LLM tag →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

LLM benchmark: Pelicans on bikes show rapid progress over two years

How we ranked this

Signal score
30 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Tool
The item describes a novel benchmark for evaluating LLMs and tracks their progress over time, including specific model performance and evolution. [lever_c_demoted from research: ic=1 ai=1.0]
Source corroboration
Single-source cluster
Only one publisher covered this so far. Single-source stories can still rank when the publisher is high-authority, but they lack cross-source corroboration.
Topics
model release, product
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
Breaking (< 6h)
Fresh story with cross-source coverage still developing. Ranking may shift as more sources report.

Full methodology in our editorial standards.

COVERAGE [1]

  1. dev.to — LLM tag TIER_1 English(EN) · Zhengxin ·

    Two years of pelicans on bicycles: 103 SVGs, 66 models, and what the guy who invented the benchmark says about it

    <p>In October 2024, Simon Willison gave sixteen LLMs the same eight words:</p> <blockquote> <p>Generate an SVG of a pelican riding a bicycle.</p> </blockquote> <p>He has been running it on every new model since. Twenty-three months later there are 103 of these drawings scattered …