A new study evaluated the diagnostic accuracy of several large language models (LLMs) in classifying colorectal polyps using the PRIME dataset. Claude Opus 4 and Gemini 2.5 Pro demonstrated the highest accuracy in differentiating polyp subtypes, performing closest to expert consensus, though overall sensitivity and specificity did not meet clinical standards. Separately, a deep learning framework named PolypVision was developed for polyp classification and segmentation, achieving high performance on public datasets and offering a device-independent solution. AI
IMPACT Demonstrates LLMs' potential in medical image analysis, while also highlighting the need for further development and validation before clinical deployment.
RANK_REASON Two research papers presenting novel applications of AI in medical diagnosis and classification.
- Claude Opus 4
- Google Gemini 2.5 Pro
- GPT-4o
- GPT-5
- GPT-o3
- Narrow-band Imaging Colorectal Endoscopic (NICE)
- Paris classification
- PRIME dataset
- CVC-ClinicDB
- EfficientNetV2-M
- Kvasir-SEG
- PolypGen
- PolypVision
- UNet++
AI-generated summary · Google Gemini · from 3 sources. How we write summaries →