PulseAugur
实时 07:34:20
English(EN) Quotient Semivalues for False-Name-Resistant Data Attribution

新机制解决机器学习数据归因中的假名操纵问题

研究人员引入了商半值作为一种新颖的机制,以解决机器学习数据归因中的假名操纵问题。该方法旨在通过聚类数据并使用代表性算子来减轻数据集拆分和重复等问题,从而提供更准确的估值。所提出的机制设计为在特定条件下具有假名防范能力,即使在数据来源不完美的情况下也能提供有界的公平性和操纵损失。在DataMarket-Gym基准测试中的初步测试表明,与传统的Shapley值相比,操纵收益显著降低。 AI

影响 引入了一种新的机器学习公平数据估值方法,有望提高训练数据归因的完整性并减少操纵。

排序理由 学术论文,介绍了一种新的数据归因技术机制。[lever_c_demoted from research: ic=1 ai=1.0]

在 arXiv cs.LG 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

新机制解决机器学习数据归因中的假名操纵问题

报道来源 [1]

  1. arXiv cs.LG TIER_1 English(EN) · Florian A. D. Burnat, Brittany I. Davidson ·

    Quotient Semivalues for False-Name-Resistant Data Attribution

    arXiv:2605.07663v2 Announce Type: replace-cross Abstract: Data valuation methods allocate payments and audit training data's contribution to machine-learning pipelines; however, they often assume passive contributors. In reality, contributors can split datasets across pseudonymou…