A new research paper explores the privacy risks associated with split language models, where a client trains a model on a server without sending its raw text. The study demonstrates that an observer with access to the model's publicly released weights can reconstruct a significant portion of the client's text from the activations and gradients exchanged during training. Specifically, experiments on GPT-2 showed that adding gradients increased token recovery from 94.20% to 97.38%, and exact document recovery from 13.71% to 37.77%. The research advocates for reporting leakage at both token and document levels, and treating the data transmitted in split models as highly sensitive. AI
IMPACT Highlights significant privacy risks in split learning setups, potentially impacting how sensitive data is handled in distributed AI training.
RANK_REASON Academic paper detailing a new finding about model privacy. [lever_c_demoted from research: ic=1 ai=1.0]
Read on Hugging Face Daily Papers →
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →