HealthBench Professional
PulseAugur coverage of HealthBench Professional — every cluster mentioning HealthBench Professional across labs, papers, and developer communities, ranked by signal.
1 day(s) with sentiment data
-
New benchmark SEER-Bench tests LLM medical knowledge updating
Researchers have developed SEER-Bench, a new benchmark for evaluating how well large language models can update their medical knowledge. The benchmark uses oncology staging data and NCCN guidelines to test models under …
-
OpenAI, Moonshot, Anthropic launch flagship models; benchmarks show varied strengths
In a rapid succession of releases, OpenAI, Moonshot AI, and Anthropic have launched their latest flagship models: GPT-5.6 Sol, Kimi K3, and Claude Opus 5, respectively. While all three models offer substantial context w…
-
New AI methods boost ML reproducibility and clinical diagnostics
Researchers are developing new methods to improve the reproducibility and benchmarking of machine learning models, particularly in specialized fields like machine health intelligence and clinical diagnostics. One approa…
-
OpenAI's ChatGPT for Clinicians outperforms doctors in tests, speeds up medical AI strategy
OpenAI has launched ChatGPT for Clinicians, a specialized AI platform designed to assist medical professionals. In clinical trials, the platform achieved a score of 59.0, significantly outperforming human doctors who sc…