A new benchmark developed by researchers at the University of California, Berkeley, has revealed that leading AI models struggle with real-world applications, scoring below 25%. OpenAI's GPT-5.5 achieved the highest score with a 24% pass rate, followed closely by Anthropic's Claude Fable 5 at 22%. Other prominent models like Google Gemini, DeepSeek, and Grok scored below 16% on tasks ranging from audio processing to theoretical physics. AI
IMPACT Highlights significant limitations in current AI capabilities for real-world tasks, suggesting a gap between theoretical performance and practical application.
RANK_REASON The cluster reports on a new benchmark and its results, which is a research output from a university.
Read on Mastodon — sigmoid.social →
- Anthropic
- ClaudeFable5
- DeepSeek
- GoogleGemini
- GPT55
- Grok
- Mythos5
- OpenAI
- STANFORD
- University of California, Berkeley
- Claude Fable 5
- Google Gemini
- GPT-5.5
AI-generated summary · Google Gemini · from 2 sources. How we write summaries →