A user tested Anthropic's Claude Fable-5 model by providing it with charts containing known numerical data. The model demonstrated impressive precision in reading these charts, accurately identifying values to the decimal point. However, the testing also revealed a limitation where the model struggled with specific counting tasks, failing to correctly identify the number 23. AI
IMPACT Highlights specific strengths and weaknesses in current LLM capabilities, informing future development and use cases.
RANK_REASON User-conducted evaluation of an AI model's capabilities. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →