PulseAugur
EN
LIVE 19:51:26

New benchmark tests Claude Opus 5, Kimi k3, Grok 4.5, Gemini 3.6 Flash

A new benchmark, baba-is-harbor, has been used to evaluate several recent large language models, including Claude Opus 5, Kimi k3, Grok 4.5, and Gemini 3.6 Flash. The benchmark, which was previously shared on Reddit, aims to assess the capabilities of these models. The evaluation also touches upon the cost-effectiveness of Claude Opus 5 compared to Fable 5 and identifies the most expensive model among those tested. AI

IMPACT Provides comparative performance data for recently released LLMs, aiding developers in model selection.

RANK_REASON The cluster involves the benchmarking of multiple LLMs on a specific task, which falls under research. [lever_c_demoted from research: ic=1 ai=1.0]

Read on r/ClaudeAI →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

New benchmark tests Claude Opus 5, Kimi k3, Grok 4.5, Gemini 3.6 Flash

COVERAGE [1]

  1. r/ClaudeAI TIER_2 English(EN) · /u/pmigdal ·

    Benchmarking Claude Opus 5, Kimi K3, Grok 4.5, and Gemini 3.6 Flash on Baba Is You

    <table> <tr><td> <a href="https://www.reddit.com/r/ClaudeAI/comments/1v9wp8t/benchmarking_claude_opus_5_kimi_k3_grok_45_and/"> <img alt="Benchmarking Claude Opus 5, Kimi K3, Grok 4.5, and Gemini 3.6 Flash on Baba Is You" src="https://external-preview.redd.it/iHD43LSXh-sb4CaMPfFLK…