PulseAugur
EN
LIVE 07:29:11

New MI-CXR benchmark reveals VLM struggles with longitudinal medical reasoning

Researchers have introduced MI-CXR, a new benchmark designed to evaluate the longitudinal reasoning capabilities of vision-language models (VLMs) when analyzing sequences of chest X-rays over time. The benchmark consists of multiple-choice questions across three task families: temporal event localization, interval-wise change reasoning, and global trajectory summarization. Initial evaluations of 14 state-of-the-art VLMs revealed an average accuracy of only 29.3%, indicating significant limitations in their ability to consistently reason about disease progression across multiple patient visits. AI

IMPACT Highlights critical limitations in current vision-language models for complex temporal reasoning, potentially guiding future research in medical AI.

RANK_REASON The item describes a new academic benchmark for evaluating AI models on a specific research task. [lever_c_demoted from research: ic=1 ai=1.0]

Read on arXiv cs.CV →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

New MI-CXR benchmark reveals VLM struggles with longitudinal medical reasoning

COVERAGE [1]

  1. arXiv cs.CV TIER_1 English(EN) · Sunghwan Steve Cho, Yunseok Han, Jaeyoung Do ·

    MI-CXR: A Benchmark for Longitudinal Reasoning over Multi-Interval Chest X-rays

    arXiv:2605.15574v2 Announce Type: replace Abstract: Longitudinal chest X-ray (CXR) interpretation requires reasoning over disease evolution across multiple patient visits, yet most existing medical VQA benchmarks focus on single images or short-horizon image pairs. We introduce M…