arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

RoutePrism:追踪智能体记忆中构建顺序的影响

RoutePrism: Tracing Construction Order Effects in Agent Memory

Dong Xu, Zhangfan Yang, Jiantao Wu, Shipeng Zhang, Zexuan Zhu, Jiangqiang Li, Jun Zhang, Junkai Ji

arXiv 2609.34160首次发表:更新:

发表机构

Shenzhen University; EasternDawn; University of Nottingham Ningbo(深圳大学; 东方黎明; 宁波诺丁汉大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

RoutePrism通过双序构建记忆并追踪差异,定位智能体记忆构建顺序的影响,证明恢复被替换记录可挽回超60%准确率,并在多个数据集和模型上验证了该诊断协议的有效性。

AI 中文摘要

以不同顺序处理相同的记录可能会丢弃不同的证据,然而仅凭最终准确率无法揭示发生了什么变化或这些变化是否重要。我们提出了RoutePrism,一种诊断协议,它从相同的源池中以两种处理顺序构建两次记忆,然后追踪哪些源、编译上下文和答案发生了变化。由于记录内容、时间戳、策略和答案模型都保持不变,任何观察到的差异都被定位到记忆构建步骤。一项匹配的四条件干预测试检验了因重排序而被替换的记录是否确实携带了任务相关证据:恢复该单条记录可挽回超过60个百分点的准确率损失,而替换为同等长度的非支持记录则无此效果。我们在PersonaMem-32K(63个主要查询,29个用户)和470个LongMemEval-S问题(历史跨越38至62个会话)上评估了该协议,并在五个答案模型上重复了核心干预。幸存者选择(定义为簇保留哪条记录的选择)驱动了大多数源级变化,而不同的记忆策略(压缩、有限最近性、MemoChat式摘要、A-MEM)在源、上下文和元数据层产生了不同的失败特征。

英文摘要

Processing the same records in a different order can discard different evidence, yet endpoint accuracy alone cannot reveal what changed or whether it mattered. We introduce RoutePrism, a diagnostic protocol that builds memory twice from the same source pool in two processing orders, then traces which sources, compiled contexts, and answers differ. Because record content, timestamps, policy, and the answer model all stay fixed, any observed difference is localized to the memory construction step. A matched four-condition intervention tests whether a record displaced by reordering actually carried task-relevant evidence: restoring that single record recovers over 60 percentage points of lost accuracy, while substituting a non-supporting record of equal length does not. We evaluate the protocol on PersonaMem-32K (63 primary queries, 29 users) and 470 LongMemEval-S questions with histories spanning 38 to 62 sessions, replicating the core intervention across five answer models. Survivor selection, defined as the choice of which record a cluster retains, drives most source-level changes, while different memory policies (compaction, bounded recency, MemoChat-style summarization, A-MEM) produce distinct failure signatures at the source, context, and metadata layers.

Comments55 pages, 5 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑