A new evaluation indicates that GPT-6 Astra successfully defends against 99.99% of direct prompt injection attacks. However, it falters in 8.5% of cases involving indirect attacks through documents. Claude Opus 5 performs better on these indirect attacks, failing only 4.8% of the time. AI
IMPACT Highlights ongoing security challenges in LLMs, particularly concerning indirect attacks via documents, which could impact the development of autonomous agents.
RANK_REASON The cluster reports on an evaluation of AI model security against prompt injection attacks. [lever_c_demoted from research: ic=1 ai=1.0]
Read on Mastodon — mastodon.social →
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →