A new arXiv paper investigates how large language models (LLMs) handle private information, finding that while models internally represent privacy norms, they still leak sensitive data. Researchers discovered that LLMs encode information type, recipient, and transmission principles as distinct internal representations. Despite this awareness, a gap between representation and actual behavior leads to privacy violations. The study proposes a method called CI-parametric steering to better control LLM privacy by intervening along these specific dimensions, suggesting that improved privacy can be achieved by aligning internal representations with desired behavior. AI
IMPACT This research could lead to more reliable methods for controlling LLM privacy, reducing data leaks in sensitive applications.
RANK_REASON Academic paper published on arXiv detailing research into LLM privacy. [lever_c_demoted from research: ic=1 ai=1.0]
- arXiv
- CI-parametric steering
- Contextual Integrity (CI) theory
- Haoran Wang
- Large Language Models
- LLMs
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →