PulseAugur
EN
LIVE 09:18:27

New Fair-ASR protocol re-evaluates LLM jailbreaks, reveals ReCode attack efficiency

Researchers have introduced Fair-ASR, a new evaluation protocol for black-box jailbreak attacks on large language models. This protocol uses shared target-call budgets to provide a more equitable comparison between different attack methods, as traditional metrics like attack success rate (ASR) can be misleading due to varying attack budgets. When applied to 11 existing attacks, Fair-ASR revealed significant shifts in attack rankings and highlighted that simple methods remain competitive. The study also introduced ReCode, a budget-efficient attack that achieved high ASR on GPT-5 with significantly fewer attacker calls. AI

IMPACT Introduces a more robust method for evaluating LLM safety and efficiency, potentially influencing future red-teaming efforts.

RANK_REASON Academic paper introducing a new evaluation protocol and attack method for LLMs. [lever_c_demoted from research: ic=1 ai=1.0]

Read on arXiv cs.AI →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

New Fair-ASR protocol re-evaluates LLM jailbreaks, reveals ReCode attack efficiency

COVERAGE [1]

  1. arXiv cs.AI TIER_1 English(EN) · Zhida He, Xiaoyu Wen, Han Qi, Ziyuan Zhou, Peng Yu, Jiajia Li, Chaochao Lu, Qiaosheng Zhang ·

    Fair ASR: Re-Evaluating Black-Box Jailbreaks under Shared Target-Call Budgets

    arXiv:2608.17360v1 Announce Type: cross Abstract: Reliable jailbreak evaluation is essential for assessing LLM safety, but most existing studies rely solely on attack success rate (ASR) without accounting for its dependence on attack budgets, resulting in unfair comparisons acros…