PulseAugur
中
实时 21:36:40
English(EN) JudgeStealer: Extracting LLM Judging Capabilities across Evaluation Protocols

LLM评判者在评估任务中表现出偏见和漏洞 · 追踪4个来源

近期研究强调了大型语言模型(LLM)评判者中存在的显著偏见和漏洞,这些评判者越来越多地被用于评估AI输出。研究表明,这些评判者容易受到模型提取攻击,其评判能力可以在不同评估协议下被高精度地复制。此外,LLM评判者表现出锚定效应偏见,即先前的分数会系统性地影响后续的判断,从而损害评估的独立性。即使在被提示考虑变化或忽略元数据的情况下,这些偏见仍然存在,影响了代码评估和其他任务的可靠性。 AI

影响 突出了LLM评估系统中的关键缺陷,需要改进方法以实现可靠的AI评估。

排序理由 多篇学术论文在arXiv上发表,详细介绍了关于LLM评判者偏见和提取的研究。

在 Hugging Face Daily Papers 阅读 →

AI 生成摘要 · Google Gemini · 来自 5 个来源。 我们如何撰写摘要 →

LLM评判者在评估任务中表现出偏见和漏洞 · 追踪4个来源

本文如何被排名

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Research
多篇学术论文在arXiv上发表,详细介绍了关于LLM评判者偏见和提取的研究。
Source corroboration
5 independent sources
Strong cross-source corroboration — multiple independent publishers covered this within the clustering window.
Topics
paper, safety
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
45 days old
Aged out of breaking-news scoring windows; ranking reflects the durable signal from the full source set.
Coverage growth since scoring
+1 source(s) since last score
New sources have picked up this story since our last re-score. Score will update on the next scoring pass.

完整方法见我们的编辑标准。

报道来源 [5]

  1. arXiv cs.CL TIER_1 English(EN) · Chen Chen, Yaolin Chen, Xuehan Sun, Juan Lin, Xueluan Gong, Yuhang Zheng, Qian Wang, Kwok-Yan Lam ·

    JudgeStealer:跨越评估协议提取LLM的评判能力

    arXiv:2608.26982v1 Announce Type: new Abstract: Large language model (LLM) judges are increasingly used across various evaluation scenarios, making their judgment capabilities valuable intellectual property. However, black-box access exposes these capabilities to model extraction…

  2. Hugging Face Daily Papers TIER_1 English(EN) ·

    JudgeStealer:跨越评估协议提取大型语言模型(LLM)的评判能力

    Large language model (LLM) judges are increasingly used across various evaluation scenarios, making their judgment capabilities valuable intellectual property. However, black-box access exposes these capabilities to model extraction attacks. Existing extraction methods do not spe…

  3. arXiv cs.CL TIER_1 English(EN) · Ante Kapetanovic, Kemal Altwlkany, Andro Mercep, Tomislav Duricic, Emanuel Lacic ·

    LLM-as-a-Judge 系统中的锚定偏差:先前的分数损害评估独立性

    arXiv:2608.25869v1 Announce Type: new Abstract: Large language models (LLMs) increasingly assess generated content, giving rise to the LLM-as-a-Judge paradigm. These systems now score outputs, filter content, and gate iterative refinement in production pipelines, where each judgm…

  4. arXiv cs.AI TIER_1 English(EN) · Jianlin Chen, Wenhui Chen, Ziyao Lin, Chi Man Vong ·

    法官应知晓何为改变:LLM作为法官评估的结构效度

    arXiv:2608.24419v1 Announce Type: new Abstract: LLM-as-a-judge evaluation is usually assessed by agreement and robustness to surface perturbations, but reliability does not establish construct validity. We formalize construct validity for an evaluator as a two-dimensional profile…

  5. arXiv cs.CL TIER_1 English(EN) · Jiwon Moon, Yerin Hwang, Dongryeol Lee, Taegwan Kang, Yongil Kim, Kyomin Jung ·

    莫以代码封面定论:探索大型语言模型代码评估中的偏见

    arXiv:2505.16222v2 Announce Type: replace Abstract: With the growing use of large language models(LLMs) as evaluators, their application has expanded to code evaluation tasks, where they assess the correctness of generated code without relying on reference implementations. While …