Researchers have developed a new method to improve the efficiency of generating radiology reports from 3D CT scans using vision-language models. The study systematically evaluated different foundation vision encoders, token-reducing projectors, and instruction-tuned large language models. The findings indicate that anatomy-guided region of interest cropping is a consistently effective strategy for improving clinical accuracy, while the PerceiverResampler paired with higher-resolution features offers the strongest configuration for resolution-based improvements. AI
IMPACT Improves efficiency and accuracy of AI-driven medical report generation, potentially aiding radiologists.
RANK_REASON The item is an academic paper detailing a new method for AI model performance. [lever_c_demoted from research: ic=1 ai=1.0]
- 2D ViT Curia
- 3D ViT Primus
- CNN
- CT-RATE
- CT Volumes from 2,398 Radiology Practices in the United States: A Real-Time Indicator of the Effect of COVID-19 on Routine Care, January to September 2020
- Foundation vision encoders
- Merlin
- MLP projector
- PerceiverResampler
- vision-language model
- Vít
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →