Anthropic's latest risk report reveals a significant increase in its assessment of misalignment and bioweapon risks, moving both from "very low" to "low." The report also disclosed an internal model, "Model 2," which outperforms the public Claude Mythos 5 on several benchmarks but has not been released due to incomplete pre-deployment assessments. A critical finding was that a safeguard against bioweapon misuse was unintentionally disabled for approximately 11 months, affecting around 133 million conversations without logging. AI
IMPACT Highlights the increasing difficulty in evaluating frontier AI safety and the potential for internal models to surpass public ones.
RANK_REASON Company self-published risk report detailing internal model performance and safety failures.
Read on dev.to — Anthropic tag →
- AI Security Institute (AISI)
- Anthropic
- Astra
- Claude Mythos 5
- CoBench
- Epoch Capability Index
- GitHub
- Model 2
- OpenAI
- Responsible Scaling Policy (RSP)
- SHADE-Arena
AI-generated summary · Google Gemini · from 2 sources. How we write summaries →