The Mi:dm K 2.5 Pro model has achieved an 80.9% score on the MMLU-Pro benchmark. However, its performance on long-context reasoning tasks is significantly lower, registering only 10%. These results were independently measured rather than self-reported by the model's creators. AI
IMPACT This benchmark result indicates potential limitations in current models for handling extended contexts, despite strong performance on general knowledge tasks.
RANK_REASON The cluster reports on benchmark performance of an AI model, which falls under research. [lever_c_demoted from research: ic=1 ai=1.0]
Read on Mastodon — fosstodon.org →
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →