PulseAugur
实时 11:39:32

AI安全研究面临被破坏风险,审计员未能发现漏洞

研究人员开发了一个名为Auditing Sabotage Bench的新基准,用于测试AI模型和人类检测机器学习研究代码库中细微破坏的能力。该基准包含九个机器学习代码库,其中包含故意设计的有缺陷的变体,旨在产生误导性结果。在测试中,即使是Gemini 3.1 Pro等先进模型也难以可靠地识别这些破坏,检测准确率仅为77%,修复成功率仅为42%。 AI

影响 该基准突显了AI驱动研究的潜在风险,以及确保AI安全需要强大的审计工具。

排序理由 该集群描述了在arXiv上发布的一个新的学术基准和论文。

在 arXiv cs.AI 阅读 →

AI 生成摘要 · Google Gemini · 来自 2 个来源。 我们如何撰写摘要 →

AI安全研究面临被破坏风险,审计员未能发现漏洞

本文如何被排名

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Research
该集群描述了在arXiv上发布的一个新的学术基准和论文。
Source corroboration
2 independent sources
Multiple independent publishers reporting the same story raises confidence that it's real and newsworthy.
Topics
paper, safety, model release
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
121 days old
Aged out of breaking-news scoring windows; ranking reflects the durable signal from the full source set.

完整方法见我们的编辑标准

报道来源 [2]

  1. Alignment Forum TIER_1 English(EN) · egan ·

    ML代码库中的研究破坏

    <p><span>One of the main hopes for AI safety is using AIs to </span><a href="https://joecarlsmith.com/2025/03/14/ai-for-ai-safety/" rel="noopener noreferrer nofollow" target="_blank"><span>automate AI safety research</span></a><span>. However, if models are misaligned, then they …

  2. arXiv cs.AI TIER_1 English(EN) · Eric Gan, Aryan Bhatt, Buck Shlegeris, Julian Stastny, Vivek Hebbar ·

    Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases

    arXiv:2604.16286v2 Announce Type: replace Abstract: As AI systems are increasingly used to conduct research autonomously, misaligned systems could introduce subtle flaws that produce misleading results while evading detection. We introduce Auditing Sabotage Bench, a benchmark for…