PulseAugur
EN
LIVE 08:17:14

New benchmark reveals multimodal LLMs struggle with scientific discovery

A new benchmark called Science Edge Evaluation (SEE) has been developed to assess the capabilities of multimodal large language models (MLLMs) in complex scientific discovery tasks. Across 19 MLLMs, the highest accuracy achieved was 48.7%, with general-purpose models performing better than science-specialized ones. Even with the use of tools, accuracy only reached 52.7%, highlighting that current MLLMs struggle to manage tool-derived information within experimental evidence boundaries and cannot reliably make evidence-bounded inferences, a crucial step for real scientific discovery. AI

IMPACT Current multimodal LLMs are not yet capable of supporting complex real laboratory science or making evidence-bounded inferences, indicating a significant gap before AI can be reliably used for novel scientific discovery.

RANK_REASON The cluster describes a new academic paper introducing a benchmark for evaluating AI models. [lever_c_demoted from research: ic=1 ai=1.0]

Read on arXiv cs.AI →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

New benchmark reveals multimodal LLMs struggle with scientific discovery

COVERAGE [1]

  1. arXiv cs.AI TIER_1 English(EN) · Taolin Han, Yuchen Zhang, Jinghang Wang, Yun Wu, Wai Yuet Chiu, Zhaohai Li, Yifei Zhang, Jinxin Wang, Yuhao Zhou, Chen Zhao, Jiajia Li, Jiaxin Li, Qile Jin, Kewei Sun, Shuang Wu, Weiqi Zhai, Renquan Lv, Junchao Li, Ruodan Chen, Qingteng Chen, Zhibo Yang,… ·

    Science Edge Evaluation: SEE the Missing Step Toward Real Scientific Discovery

    arXiv:2608.06931v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly involved in scientific discovery, yet it remains unclear whether they can support complex real laboratory science. Here we introduce Science Edge Evaluation (SEE), a multimodal benchmark…