English(EN)When T2I Synthetic Data Backfires: Amplified Privacy Risks in Real-Synthetic Mix Training
新研究发现,使用合成数据进行AI训练会放大真实数据的隐私风险
作者PulseAugur 编辑部·[4 个来源]·
新研究表明,将真实数据和合成数据结合用于训练AI模型(一种称为真实-合成混合训练(RSMT)的做法)可能会无意中放大真实数据的隐私风险。研究提出了理论框架和RSMixLeak等方法来评估和缓解这些放大的隐私泄露。此外,一份实用指南探讨了如何生成具有差分隐私(DP)的合成数据,以平衡效用和强大的隐私保证,旨在提高DP合成数据方法的信任度和采用率。
AI
arXiv:2607.13541v1 Announce Type: cross Abstract: To overcome data scarcity and privacy constraints in data collection, it has become standard practice across academia and industry to augment real training data with text-to-image (T2I)-generated synthetic data, a paradigm we term…
To overcome data scarcity and privacy constraints in data collection, it has become standard practice across academia and industry to augment real training data with text-to-image (T2I)-generated synthetic data, a paradigm we term Real-Synthetic Mix-Training (RSMT). While substit…
arXiv cs.AI
TIER_1English(EN)·Qian Ma, Sarah Rajtmajer·
arXiv:2604.07486v3 Announce Type: replace-cross Abstract: Large language models (LLMs) have emerged as a powerful tool for synthetic data generation. A particularly important use case is producing synthetic replicas of private text, which requires carefully balancing privacy and …
arXiv stat.ML
TIER_1English(EN)·Natalia Ponomareva, Zheng Xu, H. Brendan McMahan, Peter Kairouz, Lucas Rosenblatt, Vincent Cohen-Addad, Crist\'obal Guzm\'an, Ryan McKenna, Galen Andrew, Alex Bie, Da Yu, Alex Kurakin, Morteza Zadimoghaddam, Sergei Vassilvitskii, Andreas Terzis·
arXiv:2512.03238v2 Announce Type: replace-cross Abstract: High quality data is needed to unlock the full potential of AI for end users. However finding new sources of such data is getting harder: most publicly-available human generated data will soon have been used. Additionally,…