New research indicates that combining real and synthetic data for training AI models, a practice known as Real-Synthetic Mix-Training (RSMT), can inadvertently amplify privacy risks for the real data. Studies propose theoretical frameworks and methods like RSMixLeak to assess and mitigate these amplified privacy leakages. Additionally, a practical guide explores generating synthetic data with differential privacy (DP) to balance utility and strong privacy guarantees, aiming to increase trust and adoption of DP synthetic data approaches. AI
IMPACT Researchers are developing methods to mitigate privacy risks associated with synthetic data generation and training, aiming to increase trust in AI systems.
RANK_REASON The cluster consists of multiple academic papers discussing privacy risks in AI data generation and training methodologies.
- Differential Privacy
- DP synthetic data
- Natalia Ponomareva
- alphaXiv
- arXiv
- DagsHub
- Hugging Face
- large-language models
- Qian Ma
- synthetic data
- CatalyzeX
- Gotit.pub
- Real-Synthetic Mix-Training
- RSMixLeak
- ScienceCast
AI-generated summary · Google Gemini · from 4 sources. How we write summaries →