发表机构
CISPA Helmholtz Center for Information Security(亥姆霍兹信息安全中心)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对LLM交付物中修订痕迹导致的意外泄露问题,提出RevLeakBench基准,分析痕迹发生与恢复率,并设计输出端过滤器以在保留必需内容的同时降低泄露风险。
AI 中文摘要
大语言模型(LLM)助手日益帮助用户为第三方接收者起草内容。在私人起草过程中,用户或模型可能会引入某个项目,随后将其删除或替换。模型可能会从预期内容中删除该项目,但在陈述编辑时再次将其揭示。我们将此类陈述称为修订痕迹。例如,在用户共享配置文件前删除密码后,模型可能会删除该密码,但留下一条评论说:“已按请求删除密码‘No****4!’。”因此,仅看到交付文件的第三方接收者可以从评论中恢复已撤回的密码。在对三个公共对话语料库的野外分析中,我们识别出26,753个修订请求,其中2,363个(8.8%)留下修订痕迹。我们在受控条件下通过引入RevLeakBench(一个包含五个场景、共100个任务的基准,设有对话轨道和智能体轨道)更深入地研究这些痕迹。我们测量痕迹出现率、撤回项目恢复率、痕迹位置以及必需内容保留率。在六个模型中,两个轨道中约一半的交付物在撤销后陈述了编辑,而仅看到交付物的读者可以从其中约13%的交付物中恢复撤回的项目。告诉模型其整个回复将被转发给接收者,仍会在36.4%的交付物中留下修订痕迹。我们比较了提示防御和交付边界,并提出了一种输出端过滤器,该过滤器在几乎不损失必需内容的情况下大幅降低恢复率。我们相信,我们的工作有助于理解和减轻LLM交互中的意外泄露。
英文摘要
Large language model (LLM) assistants increasingly help users draft content for third-party recipients. During private drafting, the user or the model may introduce an item and later remove or replace it. The model may remove the item from the intended content but reveal it again when stating the edit. We call such statements revision traces. For example, after a user removes the password before sharing a configuration file, the model may delete it but leave a comment saying, "Removed the password 'No****4!' as requested." A third-party recipient who sees only the delivered file can therefore recover the withdrawn password from the comment. In an in-the-wild analysis of three public conversation corpora, we identify 26,753 revision requests, of which 2,363 (8.8%) leave revision traces. We study them in greater depth under controlled conditions by introducing RevLeakBench, a benchmark of 100 tasks across five scenarios with a conversation track and an agent track. We measure trace occurrence, withdrawn-item recovery, trace position, and required-content retention. Across six models, about half of the deliverables in both tracks state the edit after a revocation, and a reader that sees only the deliverable can recover the withdrawn item from about 13% of them. Telling the model that its entire reply will be forwarded to the recipient still leaves revision traces in 36.4% of the deliverables. We compare prompt defenses and a delivery boundary, and propose an output-side filter that sharply reduces recovery with little loss of required content. We believe our work can benefit efforts to understand and mitigate unintended disclosure in LLM interactions.