arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

智能体记忆系统能否追踪演化状态?

Can Agent Memory Systems Track Evolving State?

Xinyi Fan, Miri Liu, Ruozhen Yang, Siru Ouyang, Jiawei Han

arXiv 2608.19652首次发表:更新:

AI 中文总结

该研究针对LLM智能体记忆系统的状态追踪能力不足问题,构建StateMemBench基准,提出StateMem方法,可显著提升当前状态准确率,且作为轻量包装器适配现有系统。

AI 中文摘要

随着基于大语言模型(LLM)的智能体被部署用于更长、更高风险的任务,其记忆系统仍存在关键缺陷。现有记忆基准大多聚焦于回忆类任务,而我们认为,有效的记忆系统必须能追踪世界的演化状态:随着事实、约束和决策在长期交互中被修正,答案必须反映当前状态,而非已被取代的状态。我们将这一能力定义为状态追踪,并在StateMemBench中实现,该基准包含234个多会话场景,覆盖两种对话长度 regime( regime 保留原专业术语)。其闭集评分标准会判断答案是否反映当前状态、已被取代的状态,或其他情况,从结构上区分状态追踪失败与其他错误。我们的分析显示,该任务对现有记忆系统、检索增强基线及长上下文基线均具挑战性。随后,我们提出StateMem,一种以状态为先的记忆方法,明确追踪取代关系与关联依赖,结果显示,在DeepSeek-V4-Flash上,它将当前状态准确率较最强同骨干基线提升1.8倍(0.205→0.363),在Qwen-3.5-9B上较最强记忆系统提升1.6倍(0.149→0.233),同时与长上下文基线表现相当。最后,我们证明该状态方法可作为轻量单调用包装器应用于现有记忆系统,在StateMemBench的6种记忆与检索骨干上,将当前状态准确率提升32至67个百分点;长度与成本匹配的对照实验显示,其中15至32个百分点的提升源于状态结构而非新增上下文。

英文摘要

As LLM-based agents are deployed for longer and higher-stakes tasks, their memory systems continue to have crucial gaps. While existing memory benchmarks focus largely on recall-shaped tasks, we argue an effective memory system must track the evolving state of the world; as facts, constraints, and decisions are revised over a long interaction, answers must reflect the current state and not a superseded one. We define this capability as state tracking and instantiate it in StateMemBench, a benchmark of 234 multi-session scenarios spanning two conversation-length regimes. Its closed-pool grading scores whether an answer reflects the current state, the superseded state, or fails otherwise, separating state-tracking failures from other errors by construction. Our analysis shows that this task is challenging for existing memory systems, retrieval-augmented baselines, and long-context baselines. We then present StateMem, a state-first memory method that explicitly tracks supersession and relational dependencies, and show it improves current-state accuracy over the strongest same-backbone baseline by 1.8x (0.205 -> 0.363) on DeepSeek-V4-Flash and over the strongest memory system by 1.6x (0.149 -> 0.233) on Qwen-3.5-9B, while remaining competitive with the long-context baselines. Finally, we show the same state approach can be applied as a lightweight single-call wrapper over existing memory systems, lifting current-state accuracy by +32 to +67 points on StateMemBench across six memory and retrieval backends. A length- and cost-matched control attributes +15 to +32 of those points to state structure rather than added context.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑