English(EN)OpenAI hopes GPT-Red can shakedown rogue models
OpenAI 的 GPT-Red AI 黑客发现漏洞能力胜过人类 · 追踪 8 个来源
作者PulseAugur 编辑部·[13 个来源]·
OpenAI 开发了一个名为 GPT-Red 的 AI 模型,旨在充当“超级黑客”,以识别和利用其他 AI 模型中的漏洞。这个自动化的红队测试系统使用自我对抗循环,GPT-Red 攻击其他模型,而其他模型则进行防御,从而提高鲁棒性。OpenAI 表示,GPT-Red 已发现新的提示注入攻击,包括“伪造思维链”漏洞,并且在识别弱点方面比人类红队测试人员更有效,成功率为 84%,而人类为 13%。
AI
<p>CSET’s Jessica Ji shared her expert insight in an article published by MIT Technology Review. The article examines how OpenAI developed GPT-Red, an AI "super-hacker" designed to automatically identify vulnerabilities in large language models and strengthen their defenses again…
MIT Technology Review
TIER_1English(EN)·Will Douglas Heaven·
OpenAI has built an LLM super-hacker called GPT-Red that it uses as a sparring partner to help its other models boost their defenses against cyberattacks. Last week the company released the latest version of its flagship LLM, GPT-5.6. OpenAI says that training it against GPT-Red …
<p>OpenAI trained GPT-Red, an internal-only attacker model, using self-play reinforcement learning against a population of defender LLMs. It beat human red-teamers 84% to 13% on a replicated indirect prompt injection arena, found a novel "Fake Chain-of-Thought" attack class, and …
While red teaming is standard practice, using humans and AI to test the security of new models is novel. Enterprises should still ensure the model they use aligns with their business and security workflows.
OpenAI GPT-Red automates red teaming with self-play OpenAI's automated red teaming system GPT-Red uses self-play to find model weaknesses like prompt injection gaps affecting every AI user. https://www. notatechguy.com/openai-gpt-red -automates-red-teaming-with-self-play/ # NotAT…
📰 Meet GPT-Red: an LLM super-hacker OpenAI built to make its models safer OpenAI has built an LLM super-hacker called GPT-Red that it uses as a sparring partner to help its other models boost their defenses against cyberattacks. Last week the company released the latest ... 📰 Sou…
🤖 GPT-Red: Unlocking Self-Improvement for Robustness Explore GPT-Red, OpenAI’s automated red teaming system that uses self-play to improve AI safety, alignment, and prompt injection robustness. 📰 Source: OpenAI News 🔗 Link: https://openai.com/index/unlocking-self-improvement-gpt-…
OpenAIs GPT-Red findet in 84% der Testszenarien Sicherheitslücken, menschliche Red-Teamer nur in 13%. Das zeigt die operative Überlegenheit von automatisiertem Self-Play-Training im Red-Teaming gegenüber manuellen Ansätzen. https:// the-decoder.de/openais-gpt-red -findet-sicherhe…
OpenAI has built an LLM super-hacker called GPT-Red that it uses as a sparring partner to help its other models boost their defenses against cyberattacks. Last week the company released the latest… # mix # openai # ai https://www. technologyreview.com/2026/07/1 5/1140514/meet-gpt…
Model GPT-Red łamie zabezpieczenia AI z 84-procentową skutecznością, deklasując ludzkich ekspertów. OpenAI wdraża automatyczne systemy obronne, by ratować stabilność swoich modeli po kryzysie Code Red. # si # ai # sztucznainteligencja # wiadomości # informacje # technologia https…
<table> <tr><td> <a href="https://www.reddit.com/r/OpenAI/comments/1uxfkju/openai_anounces_gptred_an_ai_to_hack_its_own/"> <img alt="OpenAI anounces GPT-Red - an AI to Hack Its Own Models" src="https://preview.redd.it/j69g0q390gdh1.jpeg?width=320&crop=smart&auto=webp&…