English(EN)JudgeStealer: Extracting LLM Judging Capabilities across Evaluation Protocols
LLM评判者在评估任务中表现出偏见和漏洞 · 追踪4个来源
作者PulseAugur 编辑部·[5 个来源]·
近期研究强调了大型语言模型(LLM)评判者中存在的显著偏见和漏洞,这些评判者越来越多地被用于评估AI输出。研究表明,这些评判者容易受到模型提取攻击,其评判能力可以在不同评估协议下被高精度地复制。此外,LLM评判者表现出锚定效应偏见,即先前的分数会系统性地影响后续的判断,从而损害评估的独立性。即使在被提示考虑变化或忽略元数据的情况下,这些偏见仍然存在,影响了代码评估和其他任务的可靠性。
AI
arXiv:2608.26982v1 Announce Type: new Abstract: Large language model (LLM) judges are increasingly used across various evaluation scenarios, making their judgment capabilities valuable intellectual property. However, black-box access exposes these capabilities to model extraction…
Large language model (LLM) judges are increasingly used across various evaluation scenarios, making their judgment capabilities valuable intellectual property. However, black-box access exposes these capabilities to model extraction attacks. Existing extraction methods do not spe…
arXiv:2608.25869v1 Announce Type: new Abstract: Large language models (LLMs) increasingly assess generated content, giving rise to the LLM-as-a-Judge paradigm. These systems now score outputs, filter content, and gate iterative refinement in production pipelines, where each judgm…
arXiv cs.AI
TIER_1English(EN)·Jianlin Chen, Wenhui Chen, Ziyao Lin, Chi Man Vong·
arXiv:2608.24419v1 Announce Type: new Abstract: LLM-as-a-judge evaluation is usually assessed by agreement and robustness to surface perturbations, but reliability does not establish construct validity. We formalize construct validity for an evaluator as a two-dimensional profile…
arXiv cs.CL
TIER_1English(EN)·Jiwon Moon, Yerin Hwang, Dongryeol Lee, Taegwan Kang, Yongil Kim, Kyomin Jung·
arXiv:2505.16222v2 Announce Type: replace Abstract: With the growing use of large language models(LLMs) as evaluators, their application has expanded to code evaluation tasks, where they assess the correctness of generated code without relying on reference implementations. While …