arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

美丽心灵的永恒阳光:系统性地擦除大语言模型的记忆

Eternal Sunshine of the Spotless Mind: Systematically Erasing LLM's Memories

Olga Ohrimenko

arXiv 2609.36414首次发表:更新:

发表机构

The University of Melbourne(墨尔本大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文研究大语言模型记忆删除问题,发现现有模型无法真正遗忘用户请求删除的信息,提出DeLLM框架,通过动态构建上下文和消息来源图实现高删除率并保持实用性。

AI 中文摘要

我们考虑持久化的大语言模型(LLM),它们会随时间积累与用户交互的记忆。此类LLM使用外部存储来维护记忆,并可通过查询这些存储来克服固定上下文窗口的限制。这类系统具有众多实际应用,因为它们在响应用户查询时可以借鉴所有过去的交互。在本文中,我们探讨LLM能否在用户请求时遗忘与其共享的信息。我们发现,当前的LLM无法删除此类信息——即使它们声称已经遗忘,甚至当操作在有限上下文下进行时也是如此。为此,我们提出一个新的研究方向:LLM记忆的删除。我们表明,简单地移除与用户删除请求匹配的消息是不够的,因为对话自然会产生消息依赖关系,导致信息持续存在。为了正确处理删除请求,我们提出了DeLLM框架。该框架为每个LLM查询动态构建相关上下文,并维护一个消息来源图,以确定在删除过程中必须移除哪些消息。我们的实验表明,DeLLM在保持实用性的同时实现了高删除率。

英文摘要

We consider persistent LLMs that accumulate memories of their interactions with a user over time. Such LLMs maintain memories using external storage, which they can query to overcome the limitations of a fixed context window. Such systems have numerous practical applications, as they can draw on all past interactions when responding to user queries. In this paper, we ask whether LLMs can forget information shared with them upon a user's request. We find that current LLMs fail to delete such information---even when they claim to have forgotten it and even when operating with a limited context. To this end, we consider a new direction of study: Deletion of LLM Memories. We show that naively removing messages that match a user's deletion request is insufficient, since conversations naturally introduce message dependencies that cause information to persist. To correctly handle deletion requests, we propose the DeLLM framework. It dynamically constructs relevant context for each LLM query and maintains a provenance graph of messages to determine which ones must be removed during deletion. Our experiments show that DeLLM achieves a high deletion rate while maintaining utility.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑