An experiment was conducted to determine if AI labs are optimizing their models for a specific informal benchmark: generating an SVG of a pelican riding a bicycle. The study tested seven frontier models, including GPT-5.6 Terra and Claude Sonnet 5, by generating 1,008 SVGs across various animal-vehicle combinations. The results, analyzed using an LLM judge and Gemini 3.1 Flash-Lite, suggest that labs are not specifically 'pelicanmaxxing' their models, as performance on the pelican-bicycle prompt did not significantly outperform other combinations for the tested models. AI
IMPACT Suggests that current frontier models are not specifically optimized for niche, informal benchmarks, indicating a focus on broader capabilities.
RANK_REASON The cluster discusses an experiment and analysis of AI model performance on a specific, informal benchmark, which falls under research.
- Dylan Castillo
- Pelicanmaxxing
- Taiwan AI Labs
- Claude Sonnet 5
- DeepSeek V4 Pro
- Gemini 3.5 Flash
- GLM-5.2
- GPT-5.6 Terra
- Grok 4.5
- Mastodon
- OpenRouter
- Qwen3.7-Max
- Simon Willison
AI-generated summary · Google Gemini · from 5 sources. How we write summaries →