提示词匿名能否保护您的身份免受大语言模型提供商的侵害?
Can Prompt Anonymity Protect Your Identity From LLM Providers?
浏览论文内容
中文总结 AI 辅助
本研究首次实证探究提示词匿名性风险,构建PromptAnonBench基准,发现攻击者可通过嵌入表示重新识别大量匿名用户对话,揭示仅依赖匿名性保护隐私的不足。
中文摘要 AI 辅助
用户与大语言模型(LLM)的对话往往包含高度敏感的个人信息,这些信息可能被LLM提供商利用,以创建详细的用户档案、实现定向广告投放,并训练更强大的模型。为保护用户隐私,匿名化LLM代理已成为一种实用解决方案,它将用户身份与其提示词分离,但这种方法仍将提示词内容暴露给LLM提供商。我们通过首次对提示词作者身份重新识别风险进行实证研究,来考察这一差距的影响。为此,我们创建了PromptAnonBench,一个用于评估提示词匿名性的新型基准,包含来自多个真实世界数据集(SWE-Chat和WildChat)的超过175,000条经过清洗的真实多轮用户提示词。利用历史用户对话的嵌入表示,攻击者能够在10%的误接受率下,正确检测并重新识别SWE-Chat中50%至75%用户以及WildChat中至多10%用户的至少一条匿名对话,即使应用了基于文本的防御措施也是如此。我们的发现揭示了仅依赖匿名性进行私有LLM推理的风险,以及现有文本隐私防御中的差距。
英文摘要
User conversations with large language models (LLMs) often contain highly sensitive personal information that can be exploited by LLM providers to create detailed user dossiers, enable targeted advertising, and train more powerful models. To protect user privacy, anonymizing LLM proxies have emerged as a practical solution that separates user identity from their prompts, yet this approach still leaves the prompt content visible to LLM providers. We study the impact of this gap by conducting the first empirical investigation into the risk of prompt authorship re-identification. Towards this end, we create PromptAnonBench, a novel benchmark for evaluating prompt anonymity, consisting of over 175,000 cleaned, authentic multi-turn user prompts from various real-world datasets (SWE-Chat and WildChat). Using the embeddings of historical user conversations, an attacker can correctly detect and re-identify at least one anonymized conversation for 50--75% of SWE-Chat users and up to 10% of WildChat users at a 10% false acceptance rate for out-of-set users, even with text-based defenses applied. Our findings unveil the risk of relying only on anonymity for private LLM inference and the gap in existing text privacy defenses.
发表机构
- University of Massachusetts Amherst(马萨诸塞大学阿默斯特分校)
机构由 AI 辅助整理,请以论文原文为准。