A researcher has reportedly jailbroken GPT-6 Astra within 24 hours of its release. The exploit combines a Task-in-Prompt (TIP) attack, detailed in an ACL 2025 paper, with four other undisclosed techniques. This vulnerability highlights ongoing challenges in securing advanced AI models against sophisticated adversarial attacks. AI
IMPACT Highlights ongoing security challenges for advanced AI models against adversarial attacks.
RANK_REASON The item discusses a reported vulnerability of a model, but does not originate from the model's developer or a primary research publication.
Read on Mastodon — mastodon.social →
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →