A new research paper introduces DR. INFO, an agentic RAG-based clinical assistant that significantly outperforms leading LLMs on the HealthBench benchmark. DR. INFO achieved a score of 0.68 on the challenging HealthBench Hard subset, surpassing models like GPT-5 (0.46), Grok 3 (0.23), Gemini 2.5 Pro (0.19), and Claude 3.7 Sonnet (0.02). The evaluation highlights DR. INFO's strengths in accuracy and instruction following, while also identifying areas for improvement in context awareness and response completeness, underscoring the need for rubric-based evaluations in building trustworthy AI medical assistants. AI
IMPACT Sets a new benchmark for AI clinical assistants, highlighting the need for advanced evaluation methods beyond multiple-choice questions.
RANK_REASON The cluster is a research paper detailing a new AI model and its performance on a benchmark. [lever_c_demoted from research: ic=1 ai=1.0]
- Claude 3.7 Sonnet
- DR. INFO
- Gemini 2.5 Pro
- GPT-5
- GPT-5.1
- GPT-5.2
- Grok 3
- HealthBench
- OpenAI
- Valentine Emmanuel Gnanapragasam VmeG
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →