超越直接标识符:面向注重隐私的大语言模型查询委托的概率隐私风险估计
Beyond Direct Identifiers: Probabilistic Privacy Risk Estimation for Privacy-Conscious LLM Query Delegation
浏览论文内容
中文总结 AI 辅助
该研究针对注重隐私的LLM查询委托问题,提出概率变体PCD,结合LLM驱动的k-匿名性估计,创建PUPA-SD数据集,发现PAPILLON在Llama-3.2-3B上实现最佳隐私-效用平衡,k-匿名性是有用辅助指标。
中文摘要 AI 辅助
近期关于用户与大语言模型(LLM)交互时的隐私保护研究,多聚焦于直接、显式标识符,即标准检测器捕捉的个人可识别信息(PII),其中一种方法是注重隐私的委托(PCD),该方法由本地LLM充当中介。然而,隐私风险不仅源于显式标识符,还包括无PII的自我披露,用户可通过准标识符特征的组合被识别。我们研究一种概率变体的PCD,通过LLM驱动的k-匿名性概率估计来增强其目标。为此,我们创建了包含带有自我披露的自然用户查询的PUPA-SD数据集。初步结果表明,在PUPA-SD上优化PAPILLON,可提升多种本地模型在未见过对话上的质量,且对Llama-3.2-3B而言,能实现最佳的隐私-效用平衡,而较小模型难以同时优化质量与隐私。我们提出k-匿名性作为解决PCD问题的有用辅助指标。
英文摘要
Recent work on protecting privacy during user-LLM interactions often focuses on direct, explicit identifiers: the personally-identifiable information (PII) captured by standard detectors. One such approach is Privacy-Conscious Delegation (PCD), where a local LLM acts as an intermediary. However, privacy risk does not stem solely from explicit identifiers but also PII-free self-disclosures, leaving users identifiable through combinations of quasi-identifying traits. We investigate a probabilistic variant of PCD, where we augment its objectives with an LLM-driven probabilistic estimation of k-anonymity. To facilitate this, we first create the PUPA-SD dataset, which contains naturalistic user queries with self-disclosure. Our preliminary results indicate that optimizing PAPILLON on PUPA-SD improves quality on unseen conversations across a variety of local models and produces the best privacy-utility balance for Llama-3.2-3B, while smaller models struggle to jointly optimize quality and privacy. We propose k-anonymity as a useful auxiliary metric for tackling PCD.
发表机构
- Columbia University(哥伦比亚大学)
机构由 AI 辅助整理,请以论文原文为准。