PulseAugur
EN
LIVE 10:40:50

New Fair-ASR protocol re-evaluates LLM jailbreaks, introduces efficient ReCode attack

Researchers have introduced Fair-ASR, a new protocol for evaluating jailbreak attacks on large language models like GPT-5. This method standardizes comparisons by using shared target-call budgets, addressing limitations of previous evaluations that relied on FLOPs or ignored budget constraints. The study found that traditional attack methods remain competitive and introduced ReCode, a novel, budget-efficient attack that achieved an 85% success rate on GPT-5 with significantly fewer attacker calls. AI

IMPACT Establishes a more rigorous standard for LLM safety evaluations, potentially leading to more robust defenses against jailbreaking.

RANK_REASON The cluster reports on a new academic paper introducing a novel evaluation protocol and attack method for LLMs.

Read on Hugging Face Daily Papers →

AI-generated summary · Google Gemini · from 2 sources. How we write summaries →

New Fair-ASR protocol re-evaluates LLM jailbreaks, introduces efficient ReCode attack

COVERAGE [2]

  1. arXiv cs.AI TIER_1 English(EN) · Zhida He, Xiaoyu Wen, Han Qi, Ziyuan Zhou, Peng Yu, Jiajia Li, Chaochao Lu, Qiaosheng Zhang ·

    Fair ASR: Re-Evaluating Black-Box Jailbreaks under Shared Target-Call Budgets

    arXiv:2608.17360v1 Announce Type: cross Abstract: Reliable jailbreak evaluation is essential for assessing LLM safety, but most existing studies rely solely on attack success rate (ASR) without accounting for its dependence on attack budgets, resulting in unfair comparisons acros…

  2. Hugging Face Daily Papers TIER_1 English(EN) ·

    Fair ASR: Re-Evaluating Black-Box Jailbreaks under Shared Target-Call Budgets

    Reliable jailbreak evaluation is essential for assessing LLM safety, but most existing studies rely solely on attack success rate (ASR) without accounting for its dependence on attack budgets, resulting in unfair comparisons across methods. Existing compute-aware evaluations redu…