A new benchmark called SciReC has been developed to evaluate the relational reasoning capabilities of multimodal large language models (MLLMs). This benchmark utilizes a deficit-based diagnostic framework (DMRA) to quantify the contributions of visual understanding, knowledge exhibition, and memory recall to identify error sources. In evaluations, Claude-4.6 outperformed GPT-5.4 on overall relational scores, achieving 73% compared to GPT-5.4's 68%. The study also found that models perform worst on astronomy-related tasks and that relational reasoning is the primary cause of errors across all tested models. AI
IMPACT Establishes a new evaluation standard for multimodal LLMs, highlighting specific weaknesses in relational reasoning and domain performance.
RANK_REASON The cluster contains a research paper introducing a new benchmark for evaluating LLMs. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →