A new study published on arXiv questions the methodology of a previous Nature Medicine paper that found ChatGPT Health to be unsafe for emergency triage. The researchers argue that the original study's "exam-style" format, which constrained output and prevented clarifying questions, led to inaccurate conclusions about the AI's capability. Their own experiments, using naturalistic patient messages and various output formats, suggest that the evaluation format, rather than the model's inherent capability, significantly influences the measured triage failure rate. AI
IMPACT Highlights the critical need for realistic evaluation formats in assessing AI safety, particularly for health applications.
RANK_REASON The cluster contains an academic paper that presents new research findings and critiques existing methodology. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →