A recent study by researchers from Stanford, Harvard, and the ARISE network, named NOHARM, evaluated the performance of several AI tools in clinical settings. The study tested OpenEvidence, Doximity's Ask, OpenAI's GPT-5.6 Sol, and Anthropic's Claude Fable 5 using 1,100 real clinical cases and physician annotations. While Doximity's Ask performed best, all tested AI systems exhibited a significant flaw: 76.6% of harmful errors were omissions, meaning the AI failed to include crucial information rather than stating incorrect facts. This highlights the ongoing challenge of ensuring AI reliability in healthcare, even as regulatory bodies and legal frameworks grapple with AI accountability. AI
IMPACT Highlights critical omission errors in medical AI, emphasizing the need for human oversight and robust regulatory frameworks.
RANK_REASON The cluster reports on a new independent benchmark study evaluating AI tools in a clinical setting, including specific AI models and their performance metrics. [lever_c_demoted from research: ic=1 ai=1.0]
- Anthropic
- ARISE network
- Claude Fable 5
- Daniel Nadler
- Doximity
- Eric Topol
- GPT-5.6 Sol
- Harvard
- OpenAI
- OpenEvidence
- Stanford
- FDA
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →