arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.08716cs.AIcs.CL

记住重要时刻:用于长期任务智能体的主动记忆智能体

Remember When It Matters: Proactive Memory Agent for Long-Horizon Agents

  • Meta AI

机构由 AI 辅助整理,请以论文原文为准。

Yifan Wu, Lizhu Zhang, Yuhang Zhou, Mingyi Wang, Bo Peng, Serena Li, Xiangjun Fan, Zhuokai Zhao

AI总结:

研究长期任务中智能体决策时行为状态衰减问题,提出用主动记忆智能体并行干预,更新记忆库并决定是否提醒。该方法在多个基准测试中提升了智能体性能,消融实验证明其优势,还训练模型实现部分迁移。

AI中文摘要:

在长期任务中,与决策相关的状态常常分散在不断扩展的轨迹中,而行动智能体必须找出这些状态并采取行动。随着轨迹增长,任务要求、环境事实、先前尝试、诊断和未完成的子目标可能会被埋没在上下文窗口中或超出其范围,无法在需要时影响决策,这种失败模式被称为“行为状态衰减”。我们将记忆视为一种主动干预机制而非被动检索。一个单独的记忆智能体与未修改的行动智能体并行运行,从最近的轨迹更新结构化记忆库,并决定是注入基于记忆的提醒还是保持沉默。该模块可即插即用于前沿行动智能体和现有智能体框架。在Terminal - Bench 2.0和$\tau^2$-Bench上,它提高了较弱和较强行动智能体的pass@1,在Terminal - Bench上提升了8.3个百分点,在$\tau^2$-Bench上提升了6.8个百分点。消融实验表明选择性干预优于被动库暴露、始终注入、仅顾问指导和一般检索。作为迈向开放权重记忆策略的早期步骤,我们使用SFT和GRPO在SETA上训练Qwen3.5 - 27B,提高了验证奖励并实现了向Terminal - Bench的部分迁移。

英文摘要:

In long-horizon tasks, decision-relevant state is often scattered across an expanding trajectory, while the action agent must surface it and act. As trajectories grow, task requirements, environment facts, prior attempts, diagnoses, and open subgoals can be buried in the context window or pushed beyond it, failing to influence decisions when needed. We call this failure mode "behavioral state decay". We study memory as an active intervention mechanism rather than passive retrieval. A separate memory agent runs alongside an unmodified action agent, updating a structured memory bank from the recent trajectory and deciding whether to inject a memory-grounded reminder or remain silent. The module is plug-and-play with frontier action agents and existing agent harnesses. Across Terminal-Bench 2.0 and $τ^2$-Bench, it improves pass@1 for both weaker and stronger action agents, with gains of +8.3 pp on Terminal-Bench and +6.8 pp on $τ^2$-Bench. Ablations show that selective intervention outperforms passive bank exposure, always-on injection, advisor-only guidance, and general retrieval. As an early step toward open-weight memory policies, we train Qwen3.5-27B on SETA using SFT and GRPO, improving validation reward and achieving partial transfer to Terminal-Bench.

↑