Researchers have introduced MI-CXR, a new benchmark designed to evaluate the longitudinal reasoning capabilities of vision-language models (VLMs) when analyzing sequences of chest X-rays over time. The benchmark consists of multiple-choice questions across three task families: temporal event localization, interval-wise change reasoning, and global trajectory summarization. Initial evaluations of 14 state-of-the-art VLMs revealed an average accuracy of only 29.3%, indicating significant limitations in their ability to consistently reason about disease progression across multiple patient visits. AI
IMPACT Highlights critical limitations in current vision-language models for complex temporal reasoning, potentially guiding future research in medical AI.
RANK_REASON The item describes a new academic benchmark for evaluating AI models on a specific research task. [lever_c_demoted from research: ic=1 ai=1.0]
- alphaXiv
- arXiv
- CatalyzeX
- Chest X-rays
- DagsHub
- Gotit.pub
- Hugging Face
- MI-CXR
- ScienceCast
- Sunghwan Steve Cho
- vision-language models
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →