arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

长时程智能体中的记忆控制信号在行动之前出现

Memory Control Signals Emerge Before Action in Long Horizon Agents

Mingxuan Wang, Guorun Yao, Fei Luo, Yinglong Guo, Chao Ning, Bo Wang, Hongyue Chen, Yanbiao Ma, Jungong Han

arXiv 2609.27286首次发表:更新:

发表机构

TierFlow Team; Gaoling School of Artificial Intelligence, Renmin University of China; Tsinghua University(TierFlow团队; 中国人民大学高瓴人工智能学院; 清华大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对长时程智能体上下文管理问题,发现记忆控制信号在行动前已编码于隐藏状态,并提出PaMER框架,结合状态引导压缩与证据检索,在WorkBuddyBench上减少上下文消耗并保持性能。

AI 中文摘要

长时程语言模型智能体持续累积交互历史,这增加了计算成本,同时使得相关信息更难保留和重用。现有的上下文管理方法主要关注如何压缩或检索历史,但在很大程度上未探讨模型本身是否在记忆操作发生之前就已经表征了这些需求。我们研究了每个智能体行动之前的隐藏状态,发现压缩和召回需求已经编码在模型的内部表示中。这些信号不能简单地用上下文长度或交互进度来解释,并且它们在模型深度上表现出不同的形成模式。我们进一步表明,大多数记忆决策信息保留在紧凑的近期上下文中,而选择性恢复的历史证据补充了近期上下文所缺失的长距离依赖。基于这些发现,我们提出了带有证据检索的预行动记忆(PaMER),它将状态引导的压缩与外部证据检索相结合。PaMER+进一步引入了步骤级证据选择,以仅恢复当前任务所需的历史信息。在WorkBuddyBench上的实验,跨多个上下文管理基线和模型骨干,表明我们的框架大幅减少了上下文消耗,同时保持了具有竞争力的任务性能。

英文摘要

Long horizon language model agents continuously accumulate interaction history, increasing computational cost while making relevant information harder to preserve and reuse. Existing context management methods mainly focus on how to compress or retrieve history, but largely leave open whether the model itself already represents the need for these memory operations before they occur. We study the hidden state immediately before each agent action and find that compression and recall needs are already encoded in the model's internal representations. These signals cannot be explained by simple context length or interaction progress, and they exhibit distinct formation patterns across model depth. We further show that most memory decision information is preserved in a compact recent context, while selectively restored historical evidence complements the long range dependencies that recent context misses. Based on these findings, we propose Preaction Memory with Evidence Retrieval (PaMER), which combines state guided compression with external evidence retrieval. PaMER+ further introduces step level evidence selection to recover only the historical information required by the current task. Experiments on WorkBuddyBench, across multiple context management baselines and model backbones, show that our framework substantially reduces context consumption while maintaining competitive task performance.

Comments35 pages

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑