Researchers have introduced Fair-ASR, a new protocol for evaluating jailbreak attacks on large language models like GPT-5. This method standardizes comparisons by using shared target-call budgets, addressing limitations of previous evaluations that relied on FLOPs or ignored budget constraints. The study found that traditional attack methods remain competitive and introduced ReCode, a novel, budget-efficient attack that achieved an 85% success rate on GPT-5 with significantly fewer attacker calls. AI
IMPACT Establishes a more rigorous standard for LLM safety evaluations, potentially leading to more robust defenses against jailbreaking.
RANK_REASON The cluster reports on a new academic paper introducing a novel evaluation protocol and attack method for LLMs.
Read on Hugging Face Daily Papers →
AI-generated summary · Google Gemini · from 2 sources. How we write summaries →