Over the past two years, Simon Willison has been using a unique benchmark to track the progress of large language models: generating an SVG of a pelican riding a bicycle. Initially, models struggled with the task, producing abstract shapes and nonsensical interpretations. However, by mid-2025, models like DeepSeek R1 and Nvidia's offerings showed significant improvement, with some even beginning to "editorialise" by refusing the prompt or adding creative elements. A tournament using GPT-4.1 mini as a judge revealed Gemini 2.5 Pro preview as a top performer, highlighting the rapid advancements in LLM capabilities for code generation and image synthesis. AI
IMPACT Demonstrates rapid LLM progress in code generation and nuanced task interpretation, potentially influencing future benchmark design.
RANK_REASON The item describes a novel benchmark for evaluating LLMs and tracks their progress over time, including specific model performance and evolution. [lever_c_demoted from research: ic=1 ai=1.0]
- Claude 3.5 Sonnet
- DeepSeek-R1
- Gemini 2.5 Pro
- Gemini 3
- GPT-4.1 mini
- Llama 3.3-70B
- Mistral Small 3
- Nvidia
- Qwen2.5 Coder
- Qwen3-4B-Thinking
- Simon Willison
- SVG
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →