PulseAugur
实时 15:10:17
English(EN) Introducing GPT-Red

OpenAI 发布 GPT-Red 用于规模化 AI 安全测试 · 已追踪 5 个来源

OpenAI 推出了 GPT-Red,一个内部自动化系统,旨在大规模识别其 AI 模型中的提示注入漏洞。该系统通过对抗性自我博弈进行学习,在此过程中它会尝试利用防御模型,并利用成功的攻击来改进未来的模型。OpenAI 报告称,针对 GPT-Red 的训练已使其 GPT-5.6 Sol 模型具有显著的更强韧性,在面对提示注入攻击时,其故障率比之前的生产模型低六倍。 AI

影响 增强了 AI 模型的安全性和韧性,有可能加速更强大的 AI 系统的开发和部署。

排序理由 OpenAI 宣布了一个新的内部系统 (GPT-Red),用于提高 AI 模型的安全性和韧性,直接影响未来的模型开发。

在 X — OpenAI 阅读 →

AI 生成摘要 · Google Gemini · 来自 6 个来源。 我们如何撰写摘要 →

OpenAI 发布 GPT-Red 用于规模化 AI 安全测试 · 已追踪 5 个来源

本文如何被排名

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Frontier Release
OpenAI 宣布了一个新的内部系统 (GPT-Red),用于提高 AI 模型的安全性和韧性,直接影响未来的模型开发。
Source corroboration
6 independent sources
Strong cross-source corroboration — multiple independent publishers covered this within the clustering window.
Topics
safety, model release
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
66 days old
Aged out of breaking-news scoring windows; ranking reflects the durable signal from the full source set.
Coverage growth since scoring
+1 source(s) since last score
New sources have picked up this story since our last re-score. Score will update on the next scoring pass.

完整方法见我们的编辑标准

报道来源 [6]

  1. X — OpenAI TIER_1 English(EN) · OpenAI ·

    人工智能代理已被用于改进我们下一代模型的性能。

    AI agents are already being used to improve the capabilities of our next-generation models. We believe with GPT-Red that we have started to unlock a similar flywheel for safety, where today's models can be used to make tomorrow's models more robust, aligned, and trustworthy.

  2. X — OpenAI TIER_1 English(EN) · OpenAI ·

    针对 GPT‑Red 的训练使 GPT‑5.6 的韧性大幅提高。为了衡量这一点,我们重放了 GPT‑Red 的一些最强攻击——我们的模型无法应对其中任何一个

    Training against GPT‑Red makes GPT‑5.6 substantially more resilient. To measure this, we replayed some of GPT‑Red’s strongest attacks—none of which our models had seen during training. GPT‑5.6 Sol proved to be our most robust model against prompt injections to date, with 6× fewer

  3. X — OpenAI TIER_1 English(EN) · OpenAI ·

    GPT‑Red 通过对抗性自我博弈进行学习,其目标是提示注入各种具有挑战性的防御模型。

    GPT‑Red learns through adversarial self-play, where its goal is to prompt inject a variety of challenging defender models. Every successful attack that GPT-Red finds is used to improve these defenders, pushing GPT‑Red to continuously find broader and more complex failures.

  4. X — OpenAI TIER_1 English(EN) · OpenAI ·

    推出 GPT-Red

    Introducing GPT-Red An internal automated red teamer on a mission to find our models’ prompt injection vulnerabilities at scale, helping us build stronger defenses before wider deployment. https://t.co/GxnmxxcpSk

  5. X — OpenAI TIER_1 English(EN) · OpenAI ·

    随着模型能力的增强,安全和对齐必须与之同步扩展。

    As model capabilities grow, safety and alignment must scale with them. Red-teaming is essential, but today’s approaches are difficult to scale, creating a critical bottleneck. GPT‑Red is one way we’re addressing it.

  6. Mastodon — sigmoid.social TIER_1 English(EN) · [email protected] ·

    GPTRed,一个自动化红队模型,旨在发现AI模型的漏洞。通过使用GPT-Red对GPT-5.6进行对抗性训练,该模型

    # GPTRed , an # automatedredteaming model, was trained to discover # vulnerabilities in # AI models. By using GPT-Red to adversarially train GPT-5.6, the model became significantly more robust to # promptinjection attacks. This approach, combined with human and third-party red-te…