PulseAugur
实时 07:10:50
(CA) GPT-5.1-Codex-Max Evaluation Results

METR 发现 GPT-5.1-Codex-Max 对人工智能研发自动化构成低风险

METR 评估了 OpenAIGPT-5.1-Codex-Max,认为它是比先前模型低风险的渐进式改进。评估侧重于人工智能研发自动化和恶意复制风险,结论是当前趋势表明这些威胁在未来六个月内不太可能显著出现。然而,METR 承认不可预见的突破或计算规模的增加可能会影响这些预测。 AI

影响 表明当前人工智能发展趋势在短期内对人工智能研发自动化和恶意复制构成低风险。

排序理由 该报告是对特定模型安全影响的评估,而非新模型发布或重大政策转变。

在 METR (Model Evaluation & Threat Research) 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

METR 发现 GPT-5.1-Codex-Max 对人工智能研发自动化构成低风险

本文如何被排名

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Research
该报告是对特定模型安全影响的评估,而非新模型发布或重大政策转变。
Source corroboration
Single-source cluster
Only one publisher covered this so far. Single-source stories can still rank when the publisher is high-authority, but they lack cross-source corroboration.
Topics
safety, model release
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
279 days old
Aged out of breaking-news scoring windows; ranking reflects the durable signal from the full source set.

完整方法见我们的编辑标准

报道来源 [1]

  1. METR (Model Evaluation & Threat Research) TIER_1 (CA) ·

    GPT-5.1-Codex-Max 评估结果

    &lt;style&gt; .caption { text-align: center; color: #555; font-size: 0.9em; font-style: italic; margin-top: -0.5em; margin-bottom: 1.5em; } &lt;/style&gt; &lt;p&gt;&lt;strong&gt;Note on independence:&lt;/strong&gt; This evaluation was conducted under a standard NDA. Due to the se…