A new benchmark called SciFigBench has been developed to evaluate vision-language models (VLMs) on their behavioral reliability when presented with incomplete or misleading scientific figures. The benchmark assesses perception, reasoning, and behavioral aspects, with over 34,000 evaluation setups derived from 250 figures. The Admittance-Resistance-Inductance (A-R-I) framework was also introduced to measure how well models acknowledge uncertainty, resist misleading information, and infer cautiously. Results indicate that while GPT-5.2 excels in description quality and reasoning, it frequently hallucinates, whereas Gemini 3.1 Pro demonstrates greater uncertainty admission and resistance to misleading data. AI
IMPACT Highlights the need for VLMs to exhibit behavioral reliability beyond mere accuracy, crucial for scientific applications.
RANK_REASON The cluster contains a new academic paper introducing a novel benchmark and framework for evaluating AI models. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →