发表机构
The University of Manchester; The University of Melbourne; The University of Edinburgh; University of Southern California; The University of Texas at Austin(曼彻斯特大学; 墨尔本大学; 爱丁堡大学; 南加州大学; 德克萨斯大学奥斯汀分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对长程对话智能体的记忆可追溯性与透明性问题,提出TrajWiki框架,通过轨迹式记忆表示与Memory Wiki中间层提升对话性能及可解释性。
AI 中文摘要
大语言模型智能体在生成连贯且符合上下文的响应方面展现出强大能力,但可靠的长程对话因缺乏可追溯、可更新且具有诊断透明性的外部记忆而受到限制。现有的记忆增强智能体通常将记忆存储为孤立记录或可覆盖状态,难以保留信息随时间的起源、演变、冲突或过时情况。我们提出TrajWiki,一种面向长程对话智能体的基于轨迹的记忆框架。TrajWiki不将记忆视为静态条目,而是将每个记忆表示为基于来源的演变轨迹,通过不可变的情节快照以及ADD、REVISE、DEPRECATE等声明级操作进行维护。为减少碎片化和检索成本,TrajWiki还引入了Memory Wiki这一持久中间层,该层将对话历史逐步编译为结构化且相互关联的维基页面,捕捉重要实体、事件、数量、主题和冲突。在推理时,查询从相关维基页面分层路由至关联的记忆轨迹,再到对应的快照和来源消息,以实现基于证据的答案合成。在LoCoMo和MedMT上开展的实验表明,TrajWiki在开源和闭源LLM骨干模型上均提升了长程对话性能,同时为记忆演变、检索失败和答案生成提供了更强的可解释性和诊断可见性。
英文摘要
Large language model agents have shown strong capabilities in generating coherent and contextually appropriate responses, yet robust long-horizon dialogue remains limited by the lack of external memory that is traceable, updatable, and diagnostically transparent. Existing memory-augmented agents often store memories as isolated records or overwritable states, making it difficult to preserve how information originates, evolves, conflicts, or becomes obsolete over time. We propose TrajWiki, a trajectory-based memory framework for long-horizon conversational agents. Instead of treating memory as static entries, TrajWiki represents each memory as a source-grounded evolution trajectory, maintained through immutable episodic snapshots and claim-level operations such as ADD, REVISE, and DEPRECATE. To reduce fragmentation and retrieval cost, TrajWiki further introduces Memory Wiki, a persistent intermediate layer that incrementally compiles dialogue history into structured and interlinked wiki pages capturing salient entities, events, quantities, topics, and conflicts. At inference time, queries are routed hierarchically from relevant wiki pages to linked memory trajectories, then to corresponding snapshots and source messages for evidence-grounded answer synthesis. Experiments on LoCoMo and MedMT show that TrajWiki improves long-horizon dialogue performance across both open-source and closed-source LLM backbones, while providing greater interpretability and diagnostic visibility into memory evolution, retrieval failures, and answer generation.