PulseAugur
EN
LIVE 00:04:50

New PhysElite benchmark reveals MLLMs struggle with Olympiad-level physics

Researchers have introduced PhysElite, a new benchmark designed to evaluate the physics problem-solving capabilities of multimodal large language models (MLLMs). This benchmark features 11,586 Olympiad-tier physics problems, complete with visual diagrams, step-by-step solutions in both Chinese and English, and final answers. Initial testing on 18 different MLLMs revealed that even the top-performing model achieved only 33.7% accuracy, highlighting significant room for improvement in expert-level physical reasoning for current AI models. The evaluation also included a step-level analysis to pinpoint specific areas where models falter in their reasoning chains. AI

IMPACT Highlights limitations in current MLLMs for complex reasoning tasks, indicating a need for advancements in AI's scientific problem-solving abilities.

RANK_REASON The cluster describes a new academic paper introducing a benchmark dataset for evaluating AI models. [lever_c_demoted from research: ic=1 ai=1.0]

Read on arXiv cs.AI →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

New PhysElite benchmark reveals MLLMs struggle with Olympiad-level physics

COVERAGE [1]

  1. arXiv cs.AI TIER_1 English(EN) · Ruoran Xu, Wending Gao, Liyunfeng Chen, Aixin Shi, Haoyu Cheng, Zixiang Fang, Yiqiang Zou, Qiufeng Wang ·

    PhysElite: How Far Are LLMs from Solving Olympiad-Level Physics Problems?

    arXiv:2608.25097v2 Announce Type: replace Abstract: Understanding how (multimodal) large language models perform on physics problems requires benchmarks that reflect the difficulty and breadth of expert-level physical reasoning. Existing physics benchmarks remain limited in the f…