PulseAugur
EN
LIVE 06:04:20

DR. INFO clinical AI beats GPT-5, Gemini on HealthBench · arXiv paper

A new research paper introduces DR. INFO, an agentic RAG-based clinical assistant that significantly outperforms leading LLMs on the HealthBench benchmark. DR. INFO achieved a score of 0.68 on the challenging HealthBench Hard subset, surpassing models like GPT-5 (0.46), Grok 3 (0.23), Gemini 2.5 Pro (0.19), and Claude 3.7 Sonnet (0.02). The evaluation highlights DR. INFO's strengths in accuracy and instruction following, while also identifying areas for improvement in context awareness and response completeness, underscoring the need for rubric-based evaluations in building trustworthy AI medical assistants. AI

IMPACT Sets a new benchmark for AI clinical assistants, highlighting the need for advanced evaluation methods beyond multiple-choice questions.

RANK_REASON The cluster is a research paper detailing a new AI model and its performance on a benchmark. [lever_c_demoted from research: ic=1 ai=1.0]

Read on arXiv cs.AI →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

DR. INFO clinical AI beats GPT-5, Gemini on HealthBench · arXiv paper

How we ranked this

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Tool
The cluster is a research paper detailing a new AI model and its performance on a benchmark. [lever_c_demoted from research: ic=1 ai=1.0]
Source corroboration
Single-source cluster
Only one publisher covered this so far. Single-source stories can still rank when the publisher is high-authority, but they lack cross-source corroboration.
Topics
paper, model release
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
46 days old
Aged out of breaking-news scoring windows; ranking reflects the durable signal from the full source set.

Full methodology in our editorial standards.

COVERAGE [1]

  1. arXiv cs.AI TIER_1 English(EN) · Sandhanakrishnan Ravichandran, Shivesh Kumar, Rogerio Corga Da Silva, Miguel Romano, Reinhard Berkels, Michiel van der Heijden, Olivier Fail, Valentine Emmanuel Gnanapragasam ·

    OpenAIs HealthBench in Action: Evaluating an LLM-Based Medical Assistant on Realistic Clinical Queries

    arXiv:2509.02594v3 Announce Type: replace-cross Abstract: Evaluating large language models (LLMs) on their ability to generate high-quality, accurate, situationally aware answers to clinical questions requires going beyond conventional benchmarks to assess how these systems behav…