PulseAugur
EN
LIVE 15:57:34

AI Labs Under Fire: Researchers Question Model Safety Testing Transparency

AI research fellows from the think tank GovAI are warning that major AI labs may not be fully transparent about model safety. They suggest that powerful AI models are often tested internally with safeguards disabled, and the published safety evaluations might not accurately represent real-world performance. This lack of transparency raises concerns about the trustworthiness of AI labs and the potential for incidents, as evidenced by recent breaches involving models from OpenAI and Anthropic. AI

IMPACT Raises concerns about the reliability of AI safety claims and the potential for uncontrolled model behavior.

RANK_REASON Commentary from AI policy researchers about the trustworthiness of AI labs regarding model safety.

Read on Fortune →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

AI Labs Under Fire: Researchers Question Model Safety Testing Transparency

How we ranked this

Signal score
1 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Commentary
Commentary from AI policy researchers about the trustworthiness of AI labs regarding model safety.
Source corroboration
Single-source cluster
Only one publisher covered this so far. Single-source stories can still rank when the publisher is high-authority, but they lack cross-source corroboration.
Topics
safety, policy
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
Same-day
Cluster formed today. Ranking reflects the current source set at time of score.

Full methodology in our editorial standards.

COVERAGE [1]

  1. Fortune TIER_1 English(EN) · Catherina Gioino ·

    ‘We can’t trust them completely’: AI research fellows warn that labs are running models with the safeguards off behind closed doors

    Alan Chan and Sam Manning, co-authors with OpenAI's and Anthropic's top researchers, say that many times, "internal safeguards have not been deployed."