A developer building a health AI assistant named Tabibu found that its safety layer, designed to prevent dangerous outputs, actually lowered its score on the HealthBench benchmark. The safety layer, which includes features like emergency escalation and refusing prescription doses, was penalized for not being as comprehensive as the bare language model. This occurred because HealthBench rewards completeness, such as providing specific dosages and diagnoses, which the safety layer intentionally avoids to enhance user safety. AI
IMPACT Safety layers in health AI may be penalized by current benchmarks, potentially leading developers to prioritize completeness over safety.
RANK_REASON The item details the results of running a specific benchmark (HealthBench) on a custom AI system, including methodology and findings. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →