arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AgentMemGate:解决对话助手记忆中的推测污染问题

AgentMemGate: Addressing Speculation Contamination in Conversational Assistant Memory

Chirag Sharma, Benjamin Fowlersmith, Karime Maamari

arXiv 2610.07707首次发表:更新:

发表机构

Distyl AI(Distyl AI)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对对话助手记忆中将未决计划误记为事实的推测污染问题,提出写入时门控AgentMemGate,分类并隔离推测,消除污染并将任务准确率从65%提升至95%。

AI 中文摘要

具有长期记忆的对话式AI助手从用户消息中提取事实存入存储,供后续对话查阅。一个陈述的计划可能作为事实进入该存储:例如,用户可能搬到西雅图,却被记录为已经住在那里。我们称此为推测污染。最终状态记忆基准未能捕捉此错误,因为它们不探测中间状态且包含很少未解决的推测。我们提出AgentMemGate,一个用于档案存储记忆的写入时门控,将提取的陈述分类为推测、已完成事件、更正或其他。推测保留在记忆之外,并附带条件管理后续的提升或删除。我们还贡献了一个多会话对话数据集,其中计划被确认、放弃或保持未解决。在我们147个会话的保留集上,Mem0和Graphiti将未解决计划断言为当前状态的比例分别为35.2%和27.3%。在核心基准上,AgentMemGate相对于相同的无门控管道消除了所有观察到的污染(最暴露的提取风格从87.5%降至零),并将任务准确率从65%提升至95%。在更难的保留集上,门控污染为3.4%至5.7%,任务准确率提升9至13个百分点。我们的分析指出字段匹配是主要剩余瓶颈:现实中的推测往往不匹配任何档案字段,因此从未到达门控。我们发布了数据集、提示词和评估代码。

英文摘要

Conversational AI assistants with long-term memory extract facts from user messages into a store consulted in later conversations. A stated plan can enter that store as fact: a user who might move to Seattle may be recorded as already living there. We call this speculation contamination. Final-state memory benchmarks miss this error because they do not probe intermediate state and include few unresolved speculations. We present AgentMemGate, a write-time gate for profile-store memory that classifies extracted statements as speculation, completed event, correction, or other. Speculations remain outside memory, with conditions governing later promotion or deletion. We also contribute a dataset of multi-session conversations in which plans are confirmed, abandoned, or left unresolved. On our 147-conversation held-out set, Mem0 and Graphiti assert unresolved plans as current state for 35.2% and 27.3% of pending plans. On the core benchmark, AgentMemGate eliminates all observed contamination relative to the identical ungated pipeline (87.5% to zero for the most exposed extraction style) and raises task accuracy from 65% to 95%. On the harder held-out set, gated contamination is 3.4% to 5.7% and task accuracy rises by 9 to 13 percentage points. Our analysis identifies field matching as the main remaining bottleneck: realistic speculations often match no profile field and never reach the gate. We release our datasets, prompts, and evaluation code.

CommentsAccepted at the PALM Workshop at NeurIPS 2026. 15 pages, 4 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑