PulseAugur
EN
LIVE 17:08:46

OpenAI's GPT-Red AI hacker finds vulnerabilities better than humans · 8 sources tracked

OpenAI has developed an AI model named GPT-Red, designed to act as a "super-hacker" to identify and exploit vulnerabilities in other AI models. This automated red-teaming system uses a self-play loop, where GPT-Red attacks other models and they, in turn, defend themselves, leading to improved robustness. OpenAI states that GPT-Red has discovered novel prompt injection attacks, including a "fake chain of thought" exploit, and has proven more effective than human red-teamers in identifying weaknesses, with an 84% success rate compared to 13% for humans. AI

IMPACT Enhances AI safety by automating the discovery of vulnerabilities, potentially leading to more robust and secure AI models.

RANK_REASON Research milestone from a frontier lab detailing a new AI safety technique.

Read on Email — Mindstream →

AI-generated summary · Google Gemini · from 13 sources. How we write summaries →

OpenAI's GPT-Red AI hacker finds vulnerabilities better than humans · 8 sources tracked

How we ranked this

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Research
Research milestone from a frontier lab detailing a new AI safety technique.
Source corroboration
13 independent sources
Strong cross-source corroboration — multiple independent publishers covered this within the clustering window.
Topics
safety, model release
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
66 days old
Aged out of breaking-news scoring windows; ranking reflects the durable signal from the full source set.
Coverage growth since scoring
+5 source(s) since last score
New sources have picked up this story since our last re-score. Score will update on the next scoring pass.

Full methodology in our editorial standards.

COVERAGE [13]

  1. CSET (Georgetown — Center for Security & Emerging Tech) TIER_1 English(EN) · Jason Ly ·

    Meet GPT-Red: an LLM super-hacker OpenAI built to make its models safer

    <p>CSET’s Jessica Ji shared her expert insight in an article published by MIT Technology Review. The article examines how OpenAI developed GPT-Red, an AI "super-hacker" designed to automatically identify vulnerabilities in large language models and strengthen their defenses again…

  2. MIT Technology Review TIER_1 English(EN) · Will Douglas Heaven ·

    Meet GPT-Red: an LLM super-hacker OpenAI built to make its models safer

    OpenAI has built an LLM super-hacker called GPT-Red that it uses as a sparring partner to help its other models boost their defenses against cyberattacks. Last week the company released the latest version of its flagship LLM, GPT-5.6. OpenAI says that training it against GPT-Red …

  3. MarkTechPost TIER_1 English(EN) · Asif Razzaq ·

    OpenAI Details GPT-Red: An Internal Automated Red-Teaming Model That Beat Human Red-Teamers 84% To 13% On Prompt Injection

    <p>OpenAI trained GPT-Red, an internal-only attacker model, using self-play reinforcement learning against a population of defender LLMs. It beat human red-teamers 84% to 13% on a replicated indirect prompt injection arena, found a novel "Fake Chain-of-Thought" attack class, and …

  4. AI Business TIER_1 English(EN) · Esther Shittu ·

    OpenAI Unveils GPT-Red to Test AI Model Safety

    While red teaming is standard practice, using humans and AI to test the security of new models is novel. Enterprises should still ensure the model they use aligns with their business and security workflows.

  5. Email — Mindstream TIER_1 English(EN) · bounces+35008234-749c-ns3evnpcff6928077d7u=kill-the-newsletter.com@em5320.mindstream.news (bounces+35008234-749c-ns3evnpcff6928077d7u=kill-the-newsletter.com@em5320.mindstream.news) ·

    OpenAI hopes GPT-Red can shakedown rogue models

    <!--[if !mso]><!--><!--<![endif]-->GPT-Red is here. What is it?<!--[if mso]><xml><o:OfficeDocumentSettings><o:AllowPNG></o:AllowPNG><o:PixelsPerInch>96</o:PixelsPerInch></o:OfficeDocumentSettings></xml><![endif]--><!--[if mso]><style type="text/css"> h1, h2, h3, h4, h5, h6 {font-…

  6. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    OpenAI GPT-Red automates red teaming with self-play OpenAI's automated red teaming system GPT-Red uses self-play to find model weaknesses like prompt injection

    OpenAI GPT-Red automates red teaming with self-play OpenAI's automated red teaming system GPT-Red uses self-play to find model weaknesses like prompt injection gaps affecting every AI user. https://www. notatechguy.com/openai-gpt-red -automates-red-teaming-with-self-play/ # NotAT…

  7. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    📰 Meet GPT-Red: an LLM super-hacker OpenAI built to make its models safer OpenAI has built an LLM super-hacker called GPT-Red that it uses as a sparring partner

    📰 Meet GPT-Red: an LLM super-hacker OpenAI built to make its models safer OpenAI has built an LLM super-hacker called GPT-Red that it uses as a sparring partner to help its other models boost their defenses against cyberattacks. Last week the company released the latest ... 📰 Sou…

  8. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    🤖 GPT-Red: Unlocking Self-Improvement for Robustness Explore GPT-Red, OpenAI’s automated red teaming system that uses self-play to improve AI safety, alignment,

    🤖 GPT-Red: Unlocking Self-Improvement for Robustness Explore GPT-Red, OpenAI’s automated red teaming system that uses self-play to improve AI safety, alignment, and prompt injection robustness. 📰 Source: OpenAI News 🔗 Link: https://openai.com/index/unlocking-self-improvement-gpt-…

  9. Mastodon — mastodon.social TIER_1 日本語(JA) · [email protected] ·

    OpenAI, internal model "GPT-Red" for automatically verifying AI vulnerabilities https://www.watch.impress.co.jp/docs/news/2125829.html # watch_impress # ChatGPT # Tech # AI

    OpenAI、AIの脆弱性を自動検証する内部用モデル「GPT-Red」 https://www. watch.impress.co.jp/docs/news/ 2125829.html # watch_impress # ChatGPT # テック # AI

  10. Mastodon — mastodon.social TIER_1 Deutsch(DE) · aisyndicate ·

    OpenAI's GPT-Red finds security vulnerabilities in 84% of test scenarios, human red teamers only in 13%. This shows the operational superiority of automated

    OpenAIs GPT-Red findet in 84% der Testszenarien Sicherheitslücken, menschliche Red-Teamer nur in 13%. Das zeigt die operative Überlegenheit von automatisiertem Self-Play-Training im Red-Teaming gegenüber manuellen Ansätzen. https:// the-decoder.de/openais-gpt-red -findet-sicherhe…

  11. Mastodon — mastodon.social TIER_1 English(EN) · [email protected] ·

    OpenAI has built an LLM super-hacker called GPT-Red that it uses as a sparring partner to help its other models boost their defenses against cyberattacks. Last

    OpenAI has built an LLM super-hacker called GPT-Red that it uses as a sparring partner to help its other models boost their defenses against cyberattacks. Last week the company released the latest… # mix # openai # ai https://www. technologyreview.com/2026/07/1 5/1140514/meet-gpt…

  12. Mastodon — mastodon.social TIER_1 Polski(PL) · aisight ·

    GPT-Red Model Breaks AI Security with 84% Efficacy, Outperforming Human Experts. OpenAI Deploys Automatic Defense Systems to Save Stabi

    Model GPT-Red łamie zabezpieczenia AI z 84-procentową skutecznością, deklasując ludzkich ekspertów. OpenAI wdraża automatyczne systemy obronne, by ratować stabilność swoich modeli po kryzysie Code Red. # si # ai # sztucznainteligencja # wiadomości # informacje # technologia https…

  13. r/OpenAI TIER_2 English(EN) · /u/etherd0t ·

    OpenAI anounces GPT-Red - an AI to Hack Its Own Models

    <table> <tr><td> <a href="https://www.reddit.com/r/OpenAI/comments/1uxfkju/openai_anounces_gptred_an_ai_to_hack_its_own/"> <img alt="OpenAI anounces GPT-Red - an AI to Hack Its Own Models" src="https://preview.redd.it/j69g0q390gdh1.jpeg?width=320&amp;crop=smart&amp;auto=webp&amp;…