PulseAugur
EN
LIVE 07:50:44

New Fair-ASR protocol re-evaluates LLM jailbreaks, introduces efficient ReCode attack

Researchers have introduced Fair-ASR, a new protocol for evaluating jailbreak attacks on large language models like GPT-5. This method standardizes comparisons by using shared target-call budgets, addressing limitations of previous evaluations that relied on FLOPs or ignored budget constraints. The study found that traditional attack methods remain competitive and introduced ReCode, a novel, budget-efficient attack that achieved an 85% success rate on GPT-5 with significantly fewer attacker calls. AI

IMPACT Establishes a more rigorous standard for LLM safety evaluations, potentially leading to more robust defenses against jailbreaking.

RANK_REASON The cluster reports on a new academic paper introducing a novel evaluation protocol and attack method for LLMs.

Read on Hugging Face Daily Papers →

AI-generated summary · Google Gemini · from 3 sources. How we write summaries →

New Fair-ASR protocol re-evaluates LLM jailbreaks, introduces efficient ReCode attack

How we ranked this

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Research
The cluster reports on a new academic paper introducing a novel evaluation protocol and attack method for LLMs.
Source corroboration
3 independent sources
Multiple independent publishers reporting the same story raises confidence that it's real and newsworthy.
Topics
paper, safety
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
52 days old
Aged out of breaking-news scoring windows; ranking reflects the durable signal from the full source set.
Coverage growth since scoring
+1 source(s) since last score
New sources have picked up this story since our last re-score. Score will update on the next scoring pass.

Full methodology in our editorial standards.

COVERAGE [3]

  1. arXiv cs.AI TIER_1 English(EN) · Rishi Rajesh Shah, Chen Henry Wu, Shashwat Saxena, Ziqian Zhong, Alexander Robey, Aditi Raghunathan ·

    Jailbreaking in the Haystack

    arXiv:2511.04707v2 Announce Type: replace-cross Abstract: Recent advances in long-context language models (LMs) have enabled million-token inputs, expanding their capabilities across complex tasks like computer-use agents. Yet, the safety implications of these extended contexts r…

  2. arXiv cs.AI TIER_1 English(EN) · Zhida He, Xiaoyu Wen, Han Qi, Ziyuan Zhou, Peng Yu, Jiajia Li, Chaochao Lu, Qiaosheng Zhang ·

    Fair ASR: Re-Evaluating Black-Box Jailbreaks under Shared Target-Call Budgets

    arXiv:2608.17360v1 Announce Type: cross Abstract: Reliable jailbreak evaluation is essential for assessing LLM safety, but most existing studies rely solely on attack success rate (ASR) without accounting for its dependence on attack budgets, resulting in unfair comparisons acros…

  3. Hugging Face Daily Papers TIER_1 English(EN) ·

    Fair ASR: Re-Evaluating Black-Box Jailbreaks under Shared Target-Call Budgets

    Reliable jailbreak evaluation is essential for assessing LLM safety, but most existing studies rely solely on attack success rate (ASR) without accounting for its dependence on attack budgets, resulting in unfair comparisons across methods. Existing compute-aware evaluations redu…