PulseAugur
中
实时 06:46:34
English(EN) Fair ASR: Re-Evaluating Black-Box Jailbreaks under Shared Target-Call Budgets

新的Fair-ASR协议重新评估LLM越狱,引入高效ReCode攻击

研究人员推出了一种名为Fair-ASR的新协议,用于评估GPT-5等大型语言模型的越狱攻击。该方法通过使用共享的目标调用预算来标准化比较,解决了先前依赖FLOPs或忽略预算限制的评估方法的局限性。研究发现,传统的攻击方法仍然具有竞争力,并引入了一种新颖、预算高效的ReCode攻击,该攻击在GPT-5上取得了85%的成功率,且攻击者调用次数显著减少。 AI

影响 为LLM安全评估建立了更严格的标准,可能导致更强大的越狱防御措施。

排序理由 该集群报道了一篇介绍LLM新评估协议和攻击方法的新学术论文。

在 Hugging Face Daily Papers 阅读 →

AI 生成摘要 · Google Gemini · 来自 3 个来源。 我们如何撰写摘要 →

新的Fair-ASR协议重新评估LLM越狱,引入高效ReCode攻击

本文如何被排名

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Research
该集群报道了一篇介绍LLM新评估协议和攻击方法的新学术论文。
Source corroboration
3 independent sources
Multiple independent publishers reporting the same story raises confidence that it's real and newsworthy.
Topics
paper, safety
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
52 days old
Aged out of breaking-news scoring windows; ranking reflects the durable signal from the full source set.
Coverage growth since scoring
+1 source(s) since last score
New sources have picked up this story since our last re-score. Score will update on the next scoring pass.

完整方法见我们的编辑标准。

报道来源 [3]

  1. arXiv cs.AI TIER_1 English(EN) · Rishi Rajesh Shah, Chen Henry Wu, Shashwat Saxena, Ziqian Zhong, Alexander Robey, Aditi Raghunathan ·

    在干草堆中越狱

    arXiv:2511.04707v2 Announce Type: replace-cross Abstract: Recent advances in long-context language models (LMs) have enabled million-token inputs, expanding their capabilities across complex tasks like computer-use agents. Yet, the safety implications of these extended contexts r…

  2. arXiv cs.AI TIER_1 English(EN) · Zhida He, Xiaoyu Wen, Han Qi, Ziyuan Zhou, Peng Yu, Jiajia Li, Chaochao Lu, Qiaosheng Zhang ·

    Fair ASR:在共享目标调用预算下重新评估黑盒越狱

    arXiv:2608.17360v1 Announce Type: cross Abstract: Reliable jailbreak evaluation is essential for assessing LLM safety, but most existing studies rely solely on attack success rate (ASR) without accounting for its dependence on attack budgets, resulting in unfair comparisons acros…

  3. Hugging Face Daily Papers TIER_1 English(EN) ·

    Fair ASR:在共享目标调用预算下重新评估黑盒越狱

    Reliable jailbreak evaluation is essential for assessing LLM safety, but most existing studies rely solely on attack success rate (ASR) without accounting for its dependence on attack budgets, resulting in unfair comparisons across methods. Existing compute-aware evaluations redu…