PulseAugur
EN
LIVE 14:55:41
Français(FR) Rapport de risque d'Anthropic : modèle secret, déficit de protection sur 133 M de conversations et évaluations en panne

Anthropic raises AI risk ratings, reveals secret model, and discloses safeguard gap · 2 sources tracked

Anthropic's latest risk report reveals a significant increase in its assessment of misalignment and bioweapon risks, moving both from "very low" to "low." The report also disclosed an internal model, "Model 2," which outperforms the public Claude Mythos 5 on several benchmarks but has not been released due to incomplete pre-deployment assessments. A critical finding was that a safeguard against bioweapon misuse was unintentionally disabled for approximately 11 months, affecting around 133 million conversations without logging. AI

IMPACT Highlights the increasing difficulty in evaluating frontier AI safety and the potential for internal models to surpass public ones.

RANK_REASON Company self-published risk report detailing internal model performance and safety failures.

Read on dev.to — Anthropic tag →

AI-generated summary · Google Gemini · from 2 sources. How we write summaries →

Anthropic raises AI risk ratings, reveals secret model, and discloses safeguard gap · 2 sources tracked

COVERAGE [2]

  1. dev.to — Anthropic tag TIER_1 Français(FR) · DrMBL ·

    Anthropic Risk Report: Secret Model, Protection Deficit on 133M Conversations, and Failing Evaluations

    <p><strong>TL;DR :</strong> Le 14 août, Anthropic a publié son deuxième rapport de risque à l'échelle de l'entreprise — et s'en est servi pour relever ses propres niveaux de risque. Le risque de désalignement et le risque lié aux armes chimiques/biologiques sont tous deux passés …

  2. dev.to — Anthropic tag TIER_1 English(EN) · DrMBL ·

    Anthropic's Risk Report: A Secret Model, a 133M-Conversation Safeguard Gap, and Evals That Stopped Working

    <p><strong>TL;DR:</strong> On August 14, Anthropic published its second company-wide Risk Report — and used it to raise its own risk ratings. Misalignment and chemical/biological-weapon risk both moved from "very low" to "low," a bioweapon safeguard was found to have sat silently…