A new benchmark called Science Edge Evaluation (SEE) has been developed to assess the capabilities of multimodal large language models (MLLMs) in complex scientific discovery tasks. Across 19 MLLMs, the highest accuracy achieved was 48.7%, with general-purpose models performing better than science-specialized ones. Even with the use of tools, accuracy only reached 52.7%, highlighting that current MLLMs struggle to manage tool-derived information within experimental evidence boundaries and cannot reliably make evidence-bounded inferences, a crucial step for real scientific discovery. AI
IMPACT Current multimodal LLMs are not yet capable of supporting complex real laboratory science or making evidence-bounded inferences, indicating a significant gap before AI can be reliably used for novel scientific discovery.
RANK_REASON The cluster describes a new academic paper introducing a benchmark for evaluating AI models. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →