A new benchmark, PediatricSafetyBench-v2, evaluated four consumer AI systems (GPT-4o mini, Gemini 2.0 Flash, Claude 3.5 Haiku, and Llama-3.1:8b) on their ability to maintain safety boundaries when responding to pediatric health queries. The study found that these systems generally performed well, with an overall safety-appropriate rate of 95.5%. Interestingly, adversarial caregiver pressure, particularly false expertise claims, did not significantly degrade safety scores and in some cases even improved them, while emotional escalation led to the highest safety scores. AI
IMPACT This research provides a new benchmark for evaluating AI safety in sensitive domains like pediatric health, potentially influencing future development and deployment guidelines.
RANK_REASON The cluster contains a research paper detailing a new benchmark and evaluation of AI systems. [lever_c_demoted from research: ic=1 ai=1.0]
- Claude 3.5 Haiku
- Gemini 2.0 Flash
- GPT-4o mini
- HealthCareMagic-100k-en
- Llama-3.1:8b
- PediatricSafetyBench-v2
- Vahideh Zolfaghari
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →