发表机构
University of Saskatchewan; University of California–Irvine; University of Wisconsin–Madison; Obra D. Tompkins High School(萨斯喀彻温大学; 加州大学欧文分校; 威斯康星大学麦迪逊分校; 奥布拉·D.汤普金斯高中)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对内存增强LLM的跨域泄漏和内存诱导谄媚问题,提出推理时按领域划分内存的结构化呈现方法,在PersistBench七个模型上平均降低8.8%跨域泄漏且保持实用性。
AI 中文摘要
对话助手越来越依赖持久的长期内存来在不同会话中实现响应个性化,但当存储的用户信息被重新引入模型上下文时,也可能在不适当或不相关的场景中影响响应。我们研究了内存增强大型语言模型(LLM)中的两种此类失效模式:跨域泄漏,即来自一个生活领域的内存会影响另一个领域的响应;内存诱导的谄媚,即存储的用户信念使模型更倾向于同意用户而非如实响应。我们应用一种简单的推理时修改,即如何将内存呈现给模型,无需修改模型或内存内容。在 PersistBench 上的七个模型中,我们将常用的全上下文格式(将内存作为非结构化列表注入)与按领域划分内存的结构化格式进行比较。这种简单的修改在保持实用性的同时,持续减少了跨域泄漏,我们的最优方法相对于基线平均将泄漏降低了8.8%。
英文摘要
Conversational assistants increasingly rely on persistent long-term memory to personalize responses across sessions. However, when stored user information is reintroduced into the model context, it can also influence responses in inappropriate or unrelated settings. We study two such failure modes in memory-augmented LLMs: cross-domain leakage, where memories from one life domain affect responses in another, and memory-induced sycophancy, where stored user beliefs make models more likely to agree with the user rather than respond truthfully. We apply a simple inference-time modification to how memories are presented to the model, without changing the model or the memory contents. Across seven models on PersistBench, we compare the commonly used all-in context format, where memories are injected as an unstructured list, with structured formats that partition memories by domain. This simple modification consistently reduces cross-domain leakage while preserving utility, with our strongest method reducing leakage by $8.8\%$ on average relative to the baseline.