A new benchmark, baba-is-harbor, has been used to evaluate several recent large language models, including Claude Opus 5, Kimi k3, Grok 4.5, and Gemini 3.6 Flash. The benchmark, which was previously shared on Reddit, aims to assess the capabilities of these models. The evaluation also touches upon the cost-effectiveness of Claude Opus 5 compared to Fable 5 and identifies the most expensive model among those tested. AI
IMPACT Provides comparative performance data for recently released LLMs, aiding developers in model selection.
RANK_REASON The cluster involves the benchmarking of multiple LLMs on a specific task, which falls under research. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →