PulseAugur
实时 10:58:45
English(EN) Do LLMs Know What Is Private Internally? Probing and Steering Contextual Privacy Norms in Large Language Model Representations

研究发现:大型语言模型内部编码隐私规范但仍会泄露数据

一篇新的arXiv论文研究了大型语言模型(LLMs)如何处理隐私信息,发现虽然模型内部表征了隐私规范,但它们仍然会泄露敏感数据。研究人员发现,LLMs将信息类型、接收者和传输原则编码为不同的内部表征。尽管存在这种意识,表征与实际行为之间的差距会导致隐私泄露。该研究提出了一种名为CI-parametric steering的方法,通过干预这些特定维度来更好地控制LLM隐私,并建议通过使内部表征与期望行为保持一致来实现隐私的改进。 AI

影响 这项研究可能带来更可靠的控制LLM隐私的方法,减少敏感应用中的数据泄露。

排序理由 发布在arXiv上的学术论文,详细介绍了对LLM隐私的研究。[lever_c_demoted from research: ic=1 ai=1.0]

在 arXiv cs.CL 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

研究发现:大型语言模型内部编码隐私规范但仍会泄露数据

报道来源 [1]

  1. arXiv cs.CL TIER_1 English(EN) · Haoran Wang, Li Xiong, Kai Shu ·

    Do LLMs Know What Is Private Internally? Probing and Steering Contextual Privacy Norms in Large Language Model Representations

    arXiv:2604.00209v2 Announce Type: replace Abstract: Large language models (LLMs) are increasingly deployed in high-stakes settings, yet they frequently violate contextual privacy by disclosing private information in situations where humans would exercise discretion. This raises a…