An independent review of Anthropic's Claude Mythos 5.1 system card has found that while the model likely does not meet the CB-2 threshold for enabling chemical or biological weapon design, the assessment relied too heavily on subjective, time-intensive human evaluations. The review expressed concern that third parties were not involved in preliminary assessments or verification of risk claims. Additionally, the review noted that automated assessments may have suffered from poor elicitation and that Anthropic has not developed new automated evaluations for this risk category since May 2026. AI
IMPACT Highlights potential gaps in AI safety evaluations and the need for more robust, third-party verified risk assessments.
RANK_REASON Independent review of a model's safety system card. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →