PulseAugur
EN
LIVE 10:48:02

New benchmark SciFigBench tests VLM reliability with misleading scientific figures

A new benchmark called SciFigBench has been developed to evaluate vision-language models (VLMs) on their behavioral reliability when presented with incomplete or misleading scientific figures. The benchmark assesses perception, reasoning, and behavioral aspects, with over 34,000 evaluation setups derived from 250 figures. The Admittance-Resistance-Inductance (A-R-I) framework was also introduced to measure how well models acknowledge uncertainty, resist misleading information, and infer cautiously. Results indicate that while GPT-5.2 excels in description quality and reasoning, it frequently hallucinates, whereas Gemini 3.1 Pro demonstrates greater uncertainty admission and resistance to misleading data. AI

IMPACT Highlights the need for VLMs to exhibit behavioral reliability beyond mere accuracy, crucial for scientific applications.

RANK_REASON The cluster contains a new academic paper introducing a novel benchmark and framework for evaluating AI models. [lever_c_demoted from research: ic=1 ai=1.0]

Read on arXiv cs.AI →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

New benchmark SciFigBench tests VLM reliability with misleading scientific figures

COVERAGE [1]

  1. arXiv cs.AI TIER_1 English(EN) · Paul Osemudiame Oamen, Owusu-Banahene Osei, Ananya Mukherjee, Christian Greisinger, Steffen Eger, Pius Onobhayedo, Wei Zhao ·

    How Do VLMs Behave When Blind or Misled? Behavioral Evaluation of VLMs on Scientific Figures

    arXiv:2608.13267v1 Announce Type: cross Abstract: Existing vision-language model (VLM) benchmarks emphasize perception and reasoning accuracy (how well VLMs describe and reason about what they see in an image), with limited attention to behavioral reliability under uncertainty (h…