PulseAugur
实时 15:10:18
English(EN) OpenAI hopes GPT-Red can shakedown rogue models

OpenAI 的 GPT-Red AI 黑客发现漏洞能力胜过人类 · 追踪 8 个来源

OpenAI 开发了一个名为 GPT-Red 的 AI 模型,旨在充当“超级黑客”,以识别和利用其他 AI 模型中的漏洞。这个自动化的红队测试系统使用自我对抗循环,GPT-Red 攻击其他模型,而其他模型则进行防御,从而提高鲁棒性。OpenAI 表示,GPT-Red 已发现新的提示注入攻击,包括“伪造思维链”漏洞,并且在识别弱点方面比人类红队测试人员更有效,成功率为 84%,而人类为 13%。 AI

影响 通过自动化漏洞发现来增强 AI 安全性,有望带来更强大、更安全的 AI 模型。

排序理由 来自前沿实验室的研究里程碑,详细介绍了新的 AI 安全技术。

在 Email — Mindstream 阅读 →

AI 生成摘要 · Google Gemini · 来自 13 个来源。 我们如何撰写摘要 →

OpenAI 的 GPT-Red AI 黑客发现漏洞能力胜过人类 · 追踪 8 个来源

本文如何被排名

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Research
来自前沿实验室的研究里程碑,详细介绍了新的 AI 安全技术。
Source corroboration
13 independent sources
Strong cross-source corroboration — multiple independent publishers covered this within the clustering window.
Topics
safety, model release
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
66 days old
Aged out of breaking-news scoring windows; ranking reflects the durable signal from the full source set.
Coverage growth since scoring
+5 source(s) since last score
New sources have picked up this story since our last re-score. Score will update on the next scoring pass.

完整方法见我们的编辑标准

报道来源 [13]

  1. CSET (Georgetown — Center for Security & Emerging Tech) TIER_1 English(EN) · Jason Ly ·

    认识 GPT-Red:OpenAI 构建的用于提升模型安全性的 LLM 超级黑客

    <p>CSET’s Jessica Ji shared her expert insight in an article published by MIT Technology Review. The article examines how OpenAI developed GPT-Red, an AI "super-hacker" designed to automatically identify vulnerabilities in large language models and strengthen their defenses again…

  2. MIT Technology Review TIER_1 English(EN) · Will Douglas Heaven ·

    认识 GPT-Red:OpenAI 构建的用于提升模型安全性的 LLM 超级黑客

    OpenAI has built an LLM super-hacker called GPT-Red that it uses as a sparring partner to help its other models boost their defenses against cyberattacks. Last week the company released the latest version of its flagship LLM, GPT-5.6. OpenAI says that training it against GPT-Red …

  3. MarkTechPost TIER_1 English(EN) · Asif Razzaq ·

    OpenAI 详解 GPT-Red:一款击败人类红队 84% 对 13% 提示注入的内部自动化红队模型

    <p>OpenAI trained GPT-Red, an internal-only attacker model, using self-play reinforcement learning against a population of defender LLMs. It beat human red-teamers 84% to 13% on a replicated indirect prompt injection arena, found a novel "Fake Chain-of-Thought" attack class, and …

  4. AI Business TIER_1 English(EN) · Esther Shittu ·

    OpenAI发布GPT-Red以测试AI模型安全性

    While red teaming is standard practice, using humans and AI to test the security of new models is novel. Enterprises should still ensure the model they use aligns with their business and security workflows.

  5. Email — Mindstream TIER_1 English(EN) · bounces+35008234-749c-ns3evnpcff6928077d7u=kill-the-newsletter.com@em5320.mindstream.news (bounces+35008234-749c-ns3evnpcff6928077d7u=kill-the-newsletter.com@em5320.mindstream.news) ·

    OpenAI 希望 GPT-Red 能够震慑违规模型

    <!--[if !mso]><!--><!--<![endif]-->GPT-Red is here. What is it?<!--[if mso]><xml><o:OfficeDocumentSettings><o:AllowPNG></o:AllowPNG><o:PixelsPerInch>96</o:PixelsPerInch></o:OfficeDocumentSettings></xml><![endif]--><!--[if mso]><style type="text/css"> h1, h2, h3, h4, h5, h6 {font-…

  6. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    OpenAI GPT-Red 通过自我对弈实现红队测试自动化 OpenAI 的自动化红队测试系统 GPT-Red 利用自我对弈来发现提示注入等模型弱点

    OpenAI GPT-Red automates red teaming with self-play OpenAI's automated red teaming system GPT-Red uses self-play to find model weaknesses like prompt injection gaps affecting every AI user. https://www. notatechguy.com/openai-gpt-red -automates-red-teaming-with-self-play/ # NotAT…

  7. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    📰 认识 GPT-Red:OpenAI 构建的用于提升模型安全性的 LLM 超级黑客 OpenAI 构建了一个名为 GPT-Red 的 LLM 超级黑客,将其用作陪练伙伴

    📰 Meet GPT-Red: an LLM super-hacker OpenAI built to make its models safer OpenAI has built an LLM super-hacker called GPT-Red that it uses as a sparring partner to help its other models boost their defenses against cyberattacks. Last week the company released the latest ... 📰 Sou…

  8. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    🤖 GPT-Red:解锁鲁棒性自我改进 探索GPT-Red,OpenAI使用自我对弈来改进AI安全、对齐的自动化红队测试系统,

    🤖 GPT-Red: Unlocking Self-Improvement for Robustness Explore GPT-Red, OpenAI’s automated red teaming system that uses self-play to improve AI safety, alignment, and prompt injection robustness. 📰 Source: OpenAI News 🔗 Link: https://openai.com/index/unlocking-self-improvement-gpt-…

  9. Mastodon — mastodon.social TIER_1 日本語(JA) · [email protected] ·

    OpenAI 内部模型“GPT-Red”用于自动验证 AI 漏洞 https://www.watch.impress.co.jp/docs/news/2125829.html # watch_impress # ChatGPT # Tech # AI

    OpenAI、AIの脆弱性を自動検証する内部用モデル「GPT-Red」 https://www. watch.impress.co.jp/docs/news/ 2125829.html # watch_impress # ChatGPT # テック # AI

  10. Mastodon — mastodon.social TIER_1 Deutsch(DE) · aisyndicate ·

    OpenAI的GPT-Red在84%的测试场景中发现安全漏洞,而人类红队仅发现13%。这表明了自动化在操作上的优越性

    OpenAIs GPT-Red findet in 84% der Testszenarien Sicherheitslücken, menschliche Red-Teamer nur in 13%. Das zeigt die operative Überlegenheit von automatisiertem Self-Play-Training im Red-Teaming gegenüber manuellen Ansätzen. https:// the-decoder.de/openais-gpt-red -findet-sicherhe…

  11. Mastodon — mastodon.social TIER_1 English(EN) · [email protected] ·

    OpenAI 构建了一个名为 GPT-Red 的 LLM 超级黑客,用它作为陪练伙伴,帮助其其他模型提升网络攻击防御能力。上

    OpenAI has built an LLM super-hacker called GPT-Red that it uses as a sparring partner to help its other models boost their defenses against cyberattacks. Last week the company released the latest… # mix # openai # ai https://www. technologyreview.com/2026/07/1 5/1140514/meet-gpt…

  12. Mastodon — mastodon.social TIER_1 Polski(PL) · aisight ·

    GPT-Red模型以84%的有效性打破AI安全,超越人类专家。OpenAI部署自动防御系统以挽救稳定

    Model GPT-Red łamie zabezpieczenia AI z 84-procentową skutecznością, deklasując ludzkich ekspertów. OpenAI wdraża automatyczne systemy obronne, by ratować stabilność swoich modeli po kryzysie Code Red. # si # ai # sztucznainteligencja # wiadomości # informacje # technologia https…

  13. r/OpenAI TIER_2 English(EN) · /u/etherd0t ·

    OpenAI发布GPT-Red——一款可自主破解自身模型的AI

    <table> <tr><td> <a href="https://www.reddit.com/r/OpenAI/comments/1uxfkju/openai_anounces_gptred_an_ai_to_hack_its_own/"> <img alt="OpenAI anounces GPT-Red - an AI to Hack Its Own Models" src="https://preview.redd.it/j69g0q390gdh1.jpeg?width=320&amp;crop=smart&amp;auto=webp&amp;…