PulseAugur
EN
LIVE 09:43:40

LLMs internally encode privacy norms but still leak data, study finds

A new arXiv paper investigates how large language models (LLMs) handle private information, finding that while models internally represent privacy norms, they still leak sensitive data. Researchers discovered that LLMs encode information type, recipient, and transmission principles as distinct internal representations. Despite this awareness, a gap between representation and actual behavior leads to privacy violations. The study proposes a method called CI-parametric steering to better control LLM privacy by intervening along these specific dimensions, suggesting that improved privacy can be achieved by aligning internal representations with desired behavior. AI

IMPACT This research could lead to more reliable methods for controlling LLM privacy, reducing data leaks in sensitive applications.

RANK_REASON Academic paper published on arXiv detailing research into LLM privacy. [lever_c_demoted from research: ic=1 ai=1.0]

Read on arXiv cs.CL →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

LLMs internally encode privacy norms but still leak data, study finds

COVERAGE [1]

  1. arXiv cs.CL TIER_1 English(EN) · Haoran Wang, Li Xiong, Kai Shu ·

    Do LLMs Know What Is Private Internally? Probing and Steering Contextual Privacy Norms in Large Language Model Representations

    arXiv:2604.00209v2 Announce Type: replace Abstract: Large language models (LLMs) are increasingly deployed in high-stakes settings, yet they frequently violate contextual privacy by disclosing private information in situations where humans would exercise discretion. This raises a…