A new research paper highlights how variations in radiologist reporting practices can significantly impact the evaluation of AI-based radiology report generation (RRG) models. The study introduces a method called ReRef to rewrite reference reports, demonstrating that changes in terminology, formatting, or detail can alter model rankings. For instance, condensing normal findings in reference reports caused one model to drop in performance while another rose, suggesting current metrics may not adequately distinguish clinical interpretation from reporting style. The researchers released a new dataset, MIMIC-CXR-Ext-ReRef, to aid future research in this area. AI
IMPACT Highlights potential flaws in current AI evaluation metrics for radiology, suggesting a need for more robust methods that decouple clinical interpretation from reporting style.
RANK_REASON Research paper detailing a new methodology and dataset for evaluating AI models. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →