PulseAugur
实时 16:51:38
English(EN) Disjoint Generation of Synthetic Data

新研究探讨用于公平性和隐私的合成数据生成

两篇研究论文探讨了以公平性和隐私为重点的合成数据生成(SDG)的新方法。第一篇论文重新审视了SDG中不同影响的概念,研究了近似和估计误差如何不成比例地影响不同群体,并提出了分组SDG模型以提高效用和公平性。第二篇论文介绍了一个不相交生成模型框架,将数据集分区进行单独生成,然后将它们组合起来而不使用通用标识符,这在保持效用的同时增强了隐私和计算可行性。 AI

影响 这些论文引入了新的合成数据生成方法,可以提高在生成数据上训练的AI模型的公平性和隐私性。

排序理由 两篇在arXiv上发表的学术论文,讨论了合成数据生成的新方法。

在 arXiv cs.LG 阅读 →

AI 生成摘要 · Google Gemini · 来自 3 个来源。 我们如何撰写摘要 →

新研究探讨用于公平性和隐私的合成数据生成

报道来源 [3]

  1. arXiv cs.LG TIER_1 English(EN) · Marc Tommasi ·

    合成数据生成中的差异化影响

    We revisit the fairness notion of disparate impact for synthetic data generation (SDG), that assesses whether the utility of generated records is the same across sensitive groups. Our approach departs from existing work on fair SDG, that address the problem of correcting for undu…

  2. Hugging Face Daily Papers TIER_1 English(EN) ·

    合成数据生成中的差异化影响

    We revisit the fairness notion of disparate impact for synthetic data generation (SDG), that assesses whether the utility of generated records is the same across sensitive groups. Our approach departs from existing work on fair SDG, that address the problem of correcting for undu…

  3. arXiv cs.LG TIER_1 English(EN) · Anton Danholt Lautrup, Muhammad Rajabinasab, Tobias Hyrup, Arthur Zimek, Peter Schneider-Kamp ·

    合成数据的分离生成

    arXiv:2507.19700v2 Announce Type: replace Abstract: We propose a new framework for generating tabular synthetic datasets via disjoint generative models. In this paradigm, a dataset is partitioned into disjoint subsets that are supplied to separate instances of generative models. …