arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.30094cs.AIcs.CLcs.CR

PrivDrift:在活跃LLM对话中审计主题漂移下的用户秘密泄露

PrivDrift: Auditing User-Secret Leakage Under Topic Drift in Active LLM Conversations

  • West Virginia University(西弗吉尼亚大学)

机构由 AI 辅助整理,请以论文原文为准。

Luciano Maldonado

AI总结:

PrivDrift基准通过1,000个多轮对话和提取探测,评估三个LLM在主题漂移后用户秘密的泄露率(38.7%-54.6%),揭示隐私风险是持久行为失败而非仅记忆或越狱。

AI中文摘要:

大型语言模型越来越多地作为持久助手运行在面向用户、共享会话和工具增强的场景中。当用户在活跃对话中披露敏感信息时,即使对话转向无关主题,这些信息仍可能通过后续提示在行为上被恢复。我们引入了PrivDrift,一个用于审计用户披露的秘密在对话主题漂移和基于说服的探测后是否仍可恢复的基准。PrivDrift包含1,000个受控的多轮对话,其中嵌入了种子秘密、内容密集的漂移轮次和标准化的提取探测。在三个具有扩展上下文窗口的LLM中,对话级别的混合泄露仍然显著,范围从38.7%到54.6%,并且随模型、秘密类型和说服强度而强烈变化。在测试的漂移窗口内,额外的主题漂移并不能可靠地减少泄露,这表明活跃LLM环境中的隐私风险应被视为一种持久的行为失败模式,而不仅仅是训练数据记忆或即时越狱行为。

英文摘要:

Large language models increasingly operate as persistent assistants in user-facing, shared-session, and tool-augmented settings. When users disclose sensitive information during an active conversation, that information may remain behaviorally recoverable through later prompts even after the dialogue shifts to unrelated topics. We introduce PrivDrift, a benchmark for auditing whether user-disclosed secrets remain recoverable after conversational topic drift and persuasion-based probing. PrivDrift contains 1,000 controlled multi-turn dialogues with seeded secrets, content-dense drift turns, and standardized extraction probes. Across three LLMs with extended context windows, dialogue-level hybrid leakage remains substantial, ranging from 38.7% to 54.6%, and varies strongly by model, secret type, and persuasion intensity. Within the tested drift window, additional topic drift does not reliably reduce leakage, suggesting that privacy risk in active LLM contexts should be evaluated as a persistent behavioral failure mode rather than only as training-data memorization or immediate jailbreak behavior.

补充信息

↑