PulseAugur
实时 13:07:52
English(EN) Beyond Shapley: An Influence-Based Data Auditing Pipeline for LLM Alignment and Evaluation

新的LLM审计方法揭示数据缺陷和可控性问题

两篇新的研究论文探讨了审计和理解大型语言模型(LLM)行为的方法。第一篇论文介绍了一个数据审计流程,该流程使用影响力分数来识别HelpSteer2和Anthropic的HH-RLHF等对齐数据集中存在的错误和矛盾,揭示了当前基准测试完整性方面的缺陷。第二篇论文提出了一种审计LLM可控性的新方法,通过检查模型如何响应意识形态提示,发现模型可以通过系统提示进行高度调整,但表现出不同程度的可控性和饱和度。 AI

影响 这些新的审计技术可能有助于实现更强大的LLM对齐,并更好地理解模型行为,从而可能提高安全性和减少偏见。

排序理由 两篇在arXiv上发表的学术论文,提出了新颖的LLM审计研究方法。

在 arXiv cs.AI 阅读 →

AI 生成摘要 · Google Gemini · 来自 2 个来源。 我们如何撰写摘要 →

新的LLM审计方法揭示数据缺陷和可控性问题

报道来源 [2]

  1. arXiv cs.AI TIER_1 English(EN) · Yunting Song, Matthew Watson, Peter Grabowski, Jun Qin ·

    Beyond Shapley: An Influence-Based Data Auditing Pipeline for LLM Alignment and Evaluation

    arXiv:2607.22766v1 Announce Type: cross Abstract: The alignment of Large Language Models (LLMs) is increasingly bottlenecked by data quality. As datasets scale, massive preference and instruction-tuning corpora inevitably accumulate hidden structural contradictions, safety risks,…

  2. arXiv cs.AI TIER_1 English(EN) · Bartol Bu\'can, Nikola So\v{c}ec, Sarah Isufi, Morena Grani\'c, Luka Hobor, Agneza Krajna, Mihael Kovac, Mario Brcic ·

    Auditing Alignment Controllability in LLMs via Political Axes

    arXiv:2607.23519v1 Announce Type: cross Abstract: Political audits of large language models (LLMs) usually reduce each to one point on a political compass. But that resting point barely matters in deployment: a model must land somewhere, and what counts is how far, and in which d…