arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

长时程语言智能体中的重尾记忆痕迹

Heavy-Tailed Memory Traces in Long-Horizon Language Agents

Xinyuan Song, Zekun Cai

arXiv 2610.00010首次发表:更新:

发表机构

Emory University; The University of Tokyo; LocationMind(埃默里大学; 东京大学; LocationMind)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出核心-尾部世界模型(CTWM),通过秩基记忆控制器分配提示预算,在保持覆盖的同时减少令牌消耗并降低尾部预测误差,证明重尾记忆痕迹可作为高效智能体世界模型的控制信号。

AI 中文摘要

长时程语言智能体日益依赖外部记忆作为冻结的世界模型,然而当前记忆系统通常仅以任务成功率或令牌成本来评判。我们认为缺失的对象是记忆使用的形态:在有限上下文和重复检索下,智能体记忆可能集中于一个小的核心,而将罕见状态留在长尾中,在那里预测误差会累积。我们通过保守的尾部审计研究这一效应,发现集中现象是可复现的但依赖于策略。随机游走智能体产生对数正态兼容的检索伪影,而语义LLM策略则产生最强的截断幂律兼容的核心-尾部痕迹。受此审计启发,我们提出核心-尾部世界模型(CTWM),一种基于秩的记忆控制器,以单一指数$\ au$分配提示预算,同时保留一个汇总的尾部。在合成图世界上,CTWM保持完整的状态和转移覆盖,相对于图记忆基线,提示令牌减少5.9%,下半部分尾部预测误差降低13.6%。同样的配对比较在ALFWorld上给出一致的令牌节省,并在LongMemEval上实现24.48%的令牌减少,同时总体准确率持平。这些结果表明,重尾记忆痕迹不仅是有限检索的诊断指标,也是令牌高效智能体世界模型的实用控制信号。

英文摘要

Long-horizon language agents increasingly rely on external memory as a frozen world model, yet current memory systems are usually judged only by task success or token cost. We argue that the missing object is the shape of memory use: under finite context and repeated retrieval, agent memory can concentrate on a small core while leaving rare states in a long tail where prediction errors accumulate. We study this effect through a conservative tail audit and find that concentration is reproducible but policy-dependent. Random-walk agents produce log-normal-compatible retrieval artifacts, whereas semantic LLM policies yield the strongest truncated-power-law-compatible core--tail traces. Motivated by this audit, we propose Core--Tail World Model (CTWM), a rank-based memory controller that allocates prompt budget with a single exponent $τ$ while retaining a summarized tail. On Synthetic Graph World, CTWM preserves full state and transition coverage, reduces prompt tokens by 5.9%, and lowers bottom-half tail prediction error by 13.6% relative to a graph-memory baseline. The same paired comparison gives consistent token savings on ALFWorld and a 24.48% token reduction on LongMemEval with aggregate accuracy parity. These results suggest that heavy-tailed memory traces are not only a diagnostic of finite retrieval, but also a practical control signal for token-efficient agent world models.

CommentsUnder Review

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑