A new study by 1Password's Off-by-1 Labs reveals that over half of AI-generated security patches fail to adequately fix vulnerabilities. Researchers tested OpenAI's ChatGPT-5.5 and Anthropic's Opus 4.8 on six recent CVEs, finding that 53.9% of generated patches were flawed, either by not resolving the issue, introducing new ones, or being fragile against specific exploits. The study suggests that AI models are better at pattern-matching towards fix-like outputs than performing true reasoning for verification, and that incorrect guidance significantly degrades patch quality. AI
IMPACT Highlights the critical need for robust verification processes for AI-generated code, especially in security-sensitive applications.
RANK_REASON Study published by a security research team on the effectiveness of AI-generated code patches. [lever_c_demoted from research: ic=1 ai=1.0]
- 1Password
- Anthropic
- Cursor
- CVE-2026-22738
- CVE-2026-31431
- CVE-2026-34197
- CVE-2026-45185
- CVE-2026-8512
- GHSA-wpqr-6v78-jr5g
- Off-by-1 Labs
- OpenAI
- Opus 4.8
- The Register
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →