PulseAugur
EN
LIVE 15:07:03

Mi:dm K 2.5 Pro shows strong MMLU-Pro performance but struggles with long context

The Mi:dm K 2.5 Pro model has achieved an 80.9% score on the MMLU-Pro benchmark. However, its performance on long-context reasoning tasks is significantly lower, registering only 10%. These results were independently measured rather than self-reported by the model's creators. AI

IMPACT This benchmark result indicates potential limitations in current models for handling extended contexts, despite strong performance on general knowledge tasks.

RANK_REASON The cluster reports on benchmark performance of an AI model, which falls under research. [lever_c_demoted from research: ic=1 ai=1.0]

Read on Mastodon — fosstodon.org →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

Mi:dm K 2.5 Pro shows strong MMLU-Pro performance but struggles with long context

COVERAGE [1]

  1. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    Mi:dm K 2.5 Pro hits 80.9% on MMLU-Pro but only 10% on long-context reasoning — independently measured, not self-reported. https:// olud.ai/leaderboard.html # L

    Mi:dm K 2.5 Pro hits 80.9% on MMLU-Pro but only 10% on long-context reasoning — independently measured, not self-reported. https:// olud.ai/leaderboard.html # LLM # Benchmarks # OpenSource # AI