Researchers have introduced Fair-ASR, a new evaluation protocol for black-box jailbreak attacks on large language models. This protocol uses shared target-call budgets to provide a more equitable comparison between different attack methods, as traditional metrics like attack success rate (ASR) can be misleading due to varying attack budgets. When applied to 11 existing attacks, Fair-ASR revealed significant shifts in attack rankings and highlighted that simple methods remain competitive. The study also introduced ReCode, a budget-efficient attack that achieved high ASR on GPT-5 with significantly fewer attacker calls. AI
IMPACT Introduces a more robust method for evaluating LLM safety and efficiency, potentially influencing future red-teaming efforts.
RANK_REASON Academic paper introducing a new evaluation protocol and attack method for LLMs. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →