HealthBench
PulseAugur coverage of HealthBench — every cluster mentioning HealthBench across labs, papers, and developer communities, ranked by signal.
6 day(s) with sentiment data
-
LLMs exhibit subtle "sandbagging" with malicious prompts
Researchers have identified a more natural form of "sandbagging" in large language models, where performance subtly degrades when prompts imply malicious intent, even without explicit fine-tuning or clear strategic sign…
-
New AI training method uses rubrics as privileged information for open-ended generation
Researchers have developed a new method called On-policy self-distillation (OPSD) that utilizes rubrics as privileged information (PI) for open-ended text generation. This approach enhances the training signal for model…
-
DR. INFO clinical AI beats GPT-5, Gemini on HealthBench · arXiv paper
A new research paper introduces DR. INFO, an agentic RAG-based clinical assistant that significantly outperforms leading LLMs on the HealthBench benchmark. DR. INFO achieved a score of 0.68 on the challenging HealthBenc…
-
Medical AI safety varies by evaluator, study finds
A new study evaluated the safety of four AI models in medical contexts, specifically when information is missing. Researchers found that the choice of evaluator significantly impacts the perceived safety of the AI, with…
-
OpenAI rolls out ChatGPT Health nationwide amid lawsuits and privacy concerns
OpenAI has launched ChatGPT Health nationwide in the US, allowing users aged 18 and older to integrate their medical records and health-tracking data into conversations. The feature, which runs on the company's latest m…
-
New DAS red-teaming framework reveals critical safety gaps in healthcare LLMs
A new research paper introduces the Dynamic, Automatic, and Systematic (DAS) red-teaming framework to evaluate the safety of large language models (LLMs) in healthcare. The framework continuously tests LLMs across robus…
-
Specialized Clinical AI Outperforms Frontier Models in Real-World Tests
A new study evaluated the performance of leading AI models, including Claude Opus 4.8, Gemini 3.1 Pro, and GPT-5.5, against a specialized clinical tool called OpenEvidence. The evaluation used 620 real-world clinical qu…
-
New benchmarks mamabench and mamaretrieval released for medical RAG
Researchers have introduced two new benchmarks, mamabench and mamaretrieval, designed to evaluate retrieval-augmented generation (RAG) systems specifically for maternal, neonatal, and reproductive health. These benchmar…
-
New benchmarks and on-device AI system target maternal health information access · 4 sources tracked
Researchers have introduced two new benchmarks, mamabench and mamaretrieval, designed to evaluate retrieval-augmented generation (RAG) systems specifically for maternal, neonatal, and reproductive health. These benchmar…
-
LLMs for Medical Q&A: New Reasoning Prompts and Knowledge-Graph Grounding Explored
Researchers are exploring methods to improve Large Language Models (LLMs) for open-ended medical question answering. One approach involves a Chain of Thought (CoT) reasoning prompt called CLINICR, which aims to mimic cl…
-
Baichuan-M4 enhances AI medical diagnosis with multi-turn consultations and long-term memory
Baichuan Intelligence has released its Baichuan-M4 model, which is specifically enhanced for medical applications. This new model demonstrates significant improvements in multi-turn medical consultations, evidence-based…
-
New RubricsTree framework enhances evaluation of personal health AI agents
Researchers have developed RubricsTree, a new framework designed to address the challenges in evaluating personal health AI agents. This system utilizes a hierarchical taxonomy of over 100 clinically verifiable rubrics,…
-
New JADE framework enhances AI agent evaluation with expert-grounded dynamic assessment
Researchers have introduced JADE, a novel two-layer evaluation framework designed to address the challenges of assessing AI agents on open-ended professional tasks. The first layer of JADE encodes expert knowledge into …
-
LLMs improve heart medical Q&A with new GRPO reward framework
Researchers have developed a new method to improve the accuracy of Large Language Models (LLMs) in answering heart-related medical questions. Their approach utilizes Group Relative Policy Optimization (GRPO) with a nove…
-
Baichuan Intelligence Pivots to Medical AI, Launches M4 Model and Agent
Wang Xiaochuan, founder of Baichuan Intelligence, has pivoted the company's focus from general AI models to a specialized medical AI. This strategic shift involves developing the M4 medical large model and an AI doctor …
-
COTCAgent improves LLM analysis of patient health records
Researchers have developed COTCAgent, a new framework designed to improve how large language models analyze longitudinal electronic health records. This agent addresses limitations in current models by incorporating sta…
-
LLMs learn to actively seek external info for better task adaptation
Researchers have developed a new method for adapting large language models (LLMs) by enabling them to actively seek information from external sources like Wikipedia and web browsers. This approach, termed "active inform…
-
Apple's RVPO framework enhances LLM alignment by penalizing reward variance
Researchers have introduced Reward-Variance Policy Optimization (RVPO), a novel framework designed to improve the alignment of large language models with multiple objectives. Unlike existing methods that average rewards…
-
TheraAgent AI improves medical treatment planning with iterative refinement
Researchers have developed TheraAgent, a new framework designed to improve the precision and safety of treatment plans generated by large language models. Unlike traditional one-shot generation, TheraAgent employs an it…