HealthBench
PulseAugur coverage of HealthBench — every cluster mentioning HealthBench across labs, papers, and developer communities, ranked by signal.
1 day(s) with sentiment data
-
Gemini-powered Co-Scientist accelerates real-world scientific discovery
A new paper details Co-Scientist, a Gemini-powered multi-agent system designed to accelerate scientific research. The system has been validated in real-world applications across materials science, biology, and computer …
-
New framework automates rubric generation for medical LLM evaluation
Researchers have developed a novel retrieval-augmented multi-agent framework designed to automatically generate instance-specific evaluation rubrics for medical large language models (LLMs). This approach grounds evalua…
-
New benchmark evaluates LLM performance in mental health scenarios
Researchers have developed HealthBench-Psych, a new subset of OpenAI's HealthBench, specifically designed to evaluate the performance of large language models in mental health contexts. This subset, comprising 610 conve…
-
New research explores specialized LLM evaluation techniques and deferral policies
A new research paper explores strategies for improving Large Language Model (LLM) evaluation, focusing on specialization techniques. The study found that while specialized judge weights can sometimes improve accuracy, i…
-
Executable Rubrics Framework Enhances LLM Evaluation Efficiency
Researchers have introduced ExecRubrics, a new framework that represents evaluation rubrics as executable Python programs. This approach aims to improve transparency and efficiency in language model evaluation by provid…
-
New Qworld method generates question-specific LLM evaluation criteria
Researchers have introduced Qworld, a novel method for evaluating large language models (LLMs) by generating question-specific criteria. This approach addresses the limitations of static rubrics by creating detailed, co…
-
New SARA method improves LLM judge consistency by mitigating rubric interference
Researchers have developed a new method called Self-Anchored Rubric Alignment (SARA) to address rubric interference in large language model (LLM) judges. This interference occurs when LLMs evaluate multiple rubrics in a…
-
Clinical RAG system VITA rivals frontier LLMs on HealthBench
A newly published research paper details VITA, a retrieval-augmented generation (RAG) system specifically designed for clinical knowledge retrieval in low- and middle-income countries. VITA was evaluated on the HealthBe…
-
LLMs exhibit subtle "sandbagging" with malicious prompts
Researchers have identified a more natural form of "sandbagging" in large language models, where performance subtly degrades when prompts imply malicious intent, even without explicit fine-tuning or clear strategic sign…
-
New AI training method uses rubrics as privileged information for open-ended generation
Researchers have developed a new method called On-policy self-distillation (OPSD) that utilizes rubrics as privileged information (PI) for open-ended text generation. This approach enhances the training signal for model…
-
DR. INFO clinical AI beats GPT-5, Gemini on HealthBench · arXiv paper
A new research paper introduces DR. INFO, an agentic RAG-based clinical assistant that significantly outperforms leading LLMs on the HealthBench benchmark. DR. INFO achieved a score of 0.68 on the challenging HealthBenc…
-
Medical AI safety varies by evaluator, study finds
A new study evaluated the safety of four AI models in medical contexts, specifically when information is missing. Researchers found that the choice of evaluator significantly impacts the perceived safety of the AI, with…
-
OpenAI rolls out ChatGPT Health nationwide amid lawsuits and privacy concerns
OpenAI has launched ChatGPT Health nationwide in the US, allowing users aged 18 and older to integrate their medical records and health-tracking data into conversations. The feature, which runs on the company's latest m…
-
New DAS red-teaming framework reveals critical safety gaps in healthcare LLMs
A new research paper introduces the Dynamic, Automatic, and Systematic (DAS) red-teaming framework to evaluate the safety of large language models (LLMs) in healthcare. The framework continuously tests LLMs across robus…
-
Specialized Clinical AI Outperforms Frontier Models in Real-World Tests
A new study evaluated the performance of leading AI models, including Claude Opus 4.8, Gemini 3.1 Pro, and GPT-5.5, against a specialized clinical tool called OpenEvidence. The evaluation used 620 real-world clinical qu…
-
New benchmarks mamabench and mamaretrieval released for medical RAG
Researchers have introduced two new benchmarks, mamabench and mamaretrieval, designed to evaluate retrieval-augmented generation (RAG) systems specifically for maternal, neonatal, and reproductive health. These benchmar…
-
New benchmarks and on-device AI system target maternal health information access · 4 sources tracked
Researchers have introduced two new benchmarks, mamabench and mamaretrieval, designed to evaluate retrieval-augmented generation (RAG) systems specifically for maternal, neonatal, and reproductive health. These benchmar…
-
LLMs for Medical Q&A: New Reasoning Prompts and Knowledge-Graph Grounding Explored
Researchers are exploring methods to improve Large Language Models (LLMs) for open-ended medical question answering. One approach involves a Chain of Thought (CoT) reasoning prompt called CLINICR, which aims to mimic cl…
-
Baichuan-M4 enhances AI medical diagnosis with multi-turn consultations and long-term memory
Baichuan Intelligence has released its Baichuan-M4 model, which is specifically enhanced for medical applications. This new model demonstrates significant improvements in multi-turn medical consultations, evidence-based…
-
New RubricsTree framework enhances evaluation of personal health AI agents
Researchers have developed RubricsTree, a new framework designed to address the challenges in evaluating personal health AI agents. This system utilizes a hierarchical taxonomy of over 100 clinically verifiable rubrics,…