Researchers have developed ClinMM-Bench, a new benchmark designed to evaluate the multi-turn multimodal diagnostic reasoning capabilities of large language models (LLMs) in complex clinical scenarios. The benchmark includes 1,089 real-world clinical cases and 3,760 medical images across eight specialties. Initial evaluations of 15 representative LLMs revealed that while proprietary models showed higher diagnostic accuracy, none achieved perfect diagnoses, and all models exhibited limitations in generating reliable diagnostic reasoning, with common failure modes including information synthesis issues and visual hallucinations. AI
IMPACT This benchmark could drive improvements in LLMs for medical applications by highlighting specific reasoning failures.
RANK_REASON The cluster contains a research paper introducing a new benchmark for evaluating AI models. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →