A new study evaluated the diagnostic accuracy of several large language models (LLMs) in classifying colorectal polyps using optical images. The research utilized the PRIME dataset and compared models like Claude Opus 4, Google Gemini 2.5 Pro, GPT-o3, GPT-4o, and GPT-5 against expert responses. While most models achieved high F1 scores for broad polyp classifications, Gemini 2.5 Pro and Claude Opus 4 showed the highest accuracy in differentiating specific polyp subtypes, though overall sensitivity and specificity did not meet clinical standards. AI
IMPACT LLMs demonstrate potential in medical image analysis, but further development is needed for clinical deployment.
RANK_REASON The cluster is a research paper detailing the performance of LLMs on a specific diagnostic task. [lever_c_demoted from research: ic=1 ai=1.0]
- Claude Opus 4
- Google Gemini 2.5 Pro
- GPT-4o
- GPT-5
- GPT-o3
- Narrow-band Imaging Colorectal Endoscopic (NICE)
- Paris classification
- PRIME dataset
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →