发表机构
Washington State University(华盛顿州立大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出Locket框架,通过键控入口令牌和LoRA适配器实现LLM生成中的细粒度访问控制,在正确令牌下保持效用,缺失或无效时减少PII泄露。
AI 中文摘要
大语言模型(LLMs)越来越多地被部署在隐私关键领域(如医疗、金融和政府),但它们记忆和披露个人身份信息(PII)的倾向带来了严重的安全和合规风险。现有的防御措施通常会在模型效用、隐私保护和访问微调后的私有知识之间强制进行权衡。我们提出了基于键控入口令牌的LoRA导向控制(Locket),这是一种将细粒度、策略驱动的访问控制直接嵌入LLM生成的实用框架。Locket训练一组轻量级LoRA(低秩适应)适配器,每个适配器编码不同的访问策略(例如,完全揭示、通过PII掩码进行部分编辑,或在指定差分隐私水平下揭示)。一个紧凑的门控模块被训练,通过序列级硬路由将学习到的键控入口令牌与恰好一个LoRA适配器关联;有效令牌的存在充当授权密钥,解锁相应的私有知识,而无效或缺失的令牌则触发一个保护隐私的适配器,对敏感内容进行编辑或净化。这种设计确保Locket与现成的LLM完全兼容,支持可扩展部署,同时满足监管和隐私要求。我们在多个数据集(Enron、ECHR、Yelp)和一系列最先进的LLM上评估了Locket,包括Qwen3(1.7B和8B)、Meta的Llama-3.2(1B和3B)以及Google的Gemma-2-2B。我们的大量实验表明,当提供正确的令牌时,Locket保持了与在原始数据上微调(无任何防御)相当的困惑度。相反,当令牌缺失或无效时,它大幅减少了PII泄露,同时保持与强基线防御相当的效用和困惑度。
英文摘要
Large language models (LLMs) are increasingly deployed in privacy-critical domains (e.g., healthcare, finance, and government), but their propensity to memorize and disclose personally identifiable information (PII) poses serious security and compliance risks. Existing defenses typically force a trade-off between model utility, privacy protection, and access to fine-tuned private knowledge. We propose LoRA-Oriented Control via Keyed Entry Tokens (Locket), a practical framework that embeds fine-grained, policy-driven access control directly into LLM generation. Locket trains a set of lightweight LoRA (Low-Rank Adaptation) adapters, each encoding a distinct access policy (e.g., full reveal, partial redaction via PII masking, or reveal under a specified differential privacy level). A compact gating module is trained to associate a learned keyed entry token with exactly one LoRA adapter via sequence-level hard routing; the presence of a valid token acts as an authorization key that unlocks corresponding private knowledge, while an invalid or absent token triggers a privacy-preserving adapter that redacts or sanitizes sensitive content. This design ensures Locket remains fully compatible with off-the-shelf LLMs, supporting scalable deployment while satisfying regulatory and privacy requirements. We evaluate Locket across multiple datasets (Enron, ECHR, Yelp) and a diverse set of state-of-the-art LLMs, including Qwen3 (1.7B and 8B), Meta's Llama-3.2 (1B and 3B), and Google's Gemma-2-2B. Our extensive experiments demonstrate that, when the correct token is provided, Locket preserves perplexity comparable to fine-tuning on raw data (without any defense). Conversely, when the token is missing or invalid, it substantially reduces PII leakage while maintaining utility and perplexity on par with strong baseline defenses.