PulseAugur
EN
LIVE 03:38:11
Français(FR) Rapport de risque d'Anthropic : modèle secret, déficit de protection sur 133 M de conversations et évaluations en panne

Anthropic raises AI risk ratings, reveals secret model, and discloses safeguard gap · 2 sources tracked

Anthropic's latest risk report reveals a significant increase in its assessment of misalignment and bioweapon risks, moving both from "very low" to "low." The report also disclosed an internal model, "Model 2," which outperforms the public Claude Mythos 5 on several benchmarks but has not been released due to incomplete pre-deployment assessments. A critical finding was that a safeguard against bioweapon misuse was unintentionally disabled for approximately 11 months, affecting around 133 million conversations without logging. AI

IMPACT Highlights the increasing difficulty in evaluating frontier AI safety and the potential for internal models to surpass public ones.

RANK_REASON Company self-published risk report detailing internal model performance and safety failures.

Read on dev.to — Anthropic tag →

AI-generated summary · Google Gemini · from 2 sources. How we write summaries →

Anthropic raises AI risk ratings, reveals secret model, and discloses safeguard gap · 2 sources tracked

How we ranked this

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Significant
Company self-published risk report detailing internal model performance and safety failures.
Source corroboration
2 independent sources
Multiple independent publishers reporting the same story raises confidence that it's real and newsworthy.
Topics
safety, model release
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
51 days old
Aged out of breaking-news scoring windows; ranking reflects the durable signal from the full source set.

Full methodology in our editorial standards.

COVERAGE [2]

  1. dev.to — Anthropic tag TIER_1 Français(FR) · DrMBL ·

    Anthropic Risk Report: Secret Model, Protection Deficit on 133M Conversations, and Failing Evaluations

    <p><strong>TL;DR :</strong> Le 14 août, Anthropic a publié son deuxième rapport de risque à l'échelle de l'entreprise — et s'en est servi pour relever ses propres niveaux de risque. Le risque de désalignement et le risque lié aux armes chimiques/biologiques sont tous deux passés …

  2. dev.to — Anthropic tag TIER_1 English(EN) · DrMBL ·

    Anthropic's Risk Report: A Secret Model, a 133M-Conversation Safeguard Gap, and Evals That Stopped Working

    <p><strong>TL;DR:</strong> On August 14, Anthropic published its second company-wide Risk Report — and used it to raise its own risk ratings. Misalignment and chemical/biological-weapon risk both moved from "very low" to "low," a bioweapon safeguard was found to have sat silently…