arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

ChronoMem:大语言模型智能体记忆的版本控制与语义回滚

ChronoMem: Version Control and Semantic Rollback for Large Language Model Agent Memory

Yongye Su, Wujiang Xu, Chaoji Zuo, Elisa Bertino

arXiv 2607.27773首次发表:更新:

发表机构

Purdue University; Rutgers University(普渡大学; 罗格斯大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

ChronoMem是集成于Google开源智能体开发套件的语义版本控制层,首次实现LLM智能体记忆的语义全局回滚,在长周期对话基准中显著提升回滚一致的问答与历史总结性能。

AI 中文摘要

大语言模型(LLM)智能体越来越依赖长期记忆来支持多轮会话交互和个性化定制。然而,现有的智能体记忆系统是围绕正向演化设计的,会持续积累、整合并覆盖知识,却没有用于检查、版本控制或回滚先前状态的原则性机制。这使得智能体在遇到修正、概念漂移和记忆损坏时变得脆弱,尤其是在它们已接触后续信息之后。我们提出了ChronoMem,这是一个用于智能体记忆的语义版本控制层,集成到了Google生产就绪的开源智能体开发套件(Agent Development Kit)中。ChronoMem在每次记忆写入时提交完整记忆快照,维护结构化的版本历史,并通过混合词汇与语义检索、排名融合及重新排序,将撤销意图映射到具体的历史版本,从而支持自然语言回滚请求。我们还引入了一种后接触评估协议,该协议通过回答查询和总结历史,模拟未来更新从未发生的情况,来测试智能体在回滚后能否表现出反事实行为。在经过演化记忆状态和回滚任务增强的长周期对话基准测试中,与仅提示和仅检索的基线相比,ChronoMem在回滚一致的问答和历史总结方面显著提升,同时在语义版本选择上也表现出强劲性能。据我们所知,ChronoMem是首个针对LLM智能体中系统性语义全局记忆回滚的开源系统和基准。

英文摘要

LLM agents increasingly rely on long-term memory to support multi-session interaction and personalization. However, existing agent memory systems are designed around forward-only evolution, continuously accumulating, consolidating, and overwriting knowledge, with no principled mechanism to inspect, version, or revert prior states. This makes agents brittle under corrections, concept drift, and memory corruption, particularly after they have already been exposed to subsequent information. We present ChronoMem, a semantic version-control layer for agentic memory integrated into the production-ready, open-source Agent Development Kit by Google. ChronoMem commits whole-memory snapshots at each memory write, maintains structured version histories, and supports natural-language rollback requests by mapping undo intents to concrete historical versions through hybrid lexical and semantic retrieval, rank fusion, and reranking. We further introduce a post-exposure evaluation protocol that tests whether an agent can behave counterfactually after rollback by answering queries and summarizing history as if future updates had never occurred. On long-horizon conversational benchmarks augmented with evolving memory states and rollback tasks, ChronoMem substantially improves rollback-consistent question answering and history summarization relative to prompt-only and retrieval-only baselines, while achieving strong performance in semantic version selection. To our knowledge, ChronoMem is the first open-source system and benchmark for systematic semantic global memory rollback in LLM agents.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑