基于注意力头干预的大语言模型个性化隐私控制
Personalized Privacy Control in LLMs via Attention Head Intervention
浏览论文内容
中文总结 AI 辅助
针对大语言模型隐私控制的用户偏好差异问题,提出个性化隐私概念及P3Bench基准,发现提示策略执行效果差,进而提出Repair方法提升隐私策略依从性
中文摘要 AI 辅助
智能体AI的兴起使大语言模型(LLMs)能够访问各类用户数据,引发了严重的隐私担忧。先前关于上下文隐私的研究探讨了LLMs是否会根据上下文相关规范调整信息披露,但即便是在同一语境下,不同用户可接受的披露边界也存在差异。为解决这一局限,我们引入个性化隐私(personalized privacy),将用户特定的披露偏好纳入隐私控制。我们还提出P3Bench(个性化隐私保护基准,Personalized Privacy Preservation Benchmark),这是一个新型基准,通过加入个性化披露策略扩展了上下文隐私策略。实验表明,基于提示的策略无法可靠执行个性化隐私策略,其中Qwen2.5-7B和Gemma3-4B的平均策略无视率分别为51.25%和74.28%。最后,为解决该问题,我们提出Repair,这是一种稳健的推理时注意力头干预方法,可调整披露行为以生成符合策略的响应。我们的方法通过减少模型不遵循给定策略的情况,显著提升了对用户特定隐私偏好的依从性。
英文摘要
The rise of agentic AI enables LLMs to access diverse user data, raising critical privacy concerns. Prior work on contextual privacy studies whether LLMs regulate information disclosure according to context-dependent norms. However, acceptable disclosure boundaries may vary across users even within the same context. To address this limitation, we introduce \textit{personalized privacy}, which incorporates user-specific disclosure preferences into privacy control. We further present P3Bench~(\textbf{P}ersonalized \textbf{P}rivacy \textbf{P}reservation \textbf{Bench}mark), a novel benchmark extending contextual privacy policies with personalized disclosure policies. Experiments show that prompt-based policies fail to reliably enforce personalized privacy policies, with Qwen2.5-7B and Gemma3-4B showing average policy ignorance ratios of 51.25\% and 74.28\%, respectively. Finally, to address this problem, we propose \textsc{Repair}, a robust inference-time attention head intervention method that adjusts disclosure behavior toward policy-consistent responses. Our method significantly improves adherence to user-specific privacy preferences by reducing cases where the model fails to follow the given policy.
发表机构
- IPAI, Seoul National University(首尔国立大学IPAI)
- Max Planck Institute for Software Systems(马克斯·普朗克软件系统研究所)
机构由 AI 辅助整理,请以论文原文为准。