发表机构
North Carolina State University; Shopify(北卡罗来纳州立大学; Shopify)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
AMBER提出追加式记忆框架,通过强化学习训练网络代理联合推理、行动和写记忆,保证信息保留,在WebArena Lite上比覆盖式记忆平均成功率提升4.09个百分点,兼顾上下文效率与长时程可靠性。
AI 中文摘要
现代语言模型代理越来越多地在长时程、多步骤轨迹中与外部环境交互,其中累积的交互历史可能迅速超出实际上下文预算。为确保可靠性,代理必须在长时程内保持事实信息,记住执行错误和纠正性反馈,并跟踪跨操作的进度。已有多种方法被提出,以避免在上下文中维护整个执行历史,例如使用推理和操作历史、通过覆盖机制学习维护固定大小的记忆,以及周期性总结。尽管覆盖式记忆原则上可以保留追加式记忆所能保留的任何内容,但它必须学会将每个事实贯穿于后续的每次重写,而这在稀疏的结果奖励下难以学习;对于网络代理等交互式应用,我们发现训练后的覆盖式记忆会删除轨迹所需的关键信息,以及从环境接收的纠正性反馈。我们引入了AMBER(用于证据保留的追加式记忆库)——一个简单且可扩展的框架,其中代理联合学习推理、行动和写入自由形式的记忆,而追加式规则通过构造保证保留。这使得AMBER能够使用结果奖励进行端到端的强化学习训练,而无需大量精心策划的SFT数据。在WebArena Lite上,AMBER将基于覆盖式记忆的平均成功率提高了4.09个百分点,将五次重复运行中解决的任务比例提高了4.8个百分点,并匹配了在更昂贵的策划监督下训练的覆盖式基线。AMBER在实现这些改进的同时保持了实用的令牌预算,在上下文效率、任务性能和可靠的长时程执行之间提供了强有力的平衡。
英文摘要
Modern language-model agents increasingly interact with external environments over long-horizon, multi-step trajectories, where the accumulated interaction history can quickly exceed practical context budgets. To ensure reliability, agents must maintain factual information over long horizons, remember execution errors and corrective feedback, and track progress across actions. Several approaches have been proposed to achieve this without the need for maintaining the entire execution history in context, such as using the reasoning and action history, learning to maintain a fixed-size memory through an overwrite mechanism, and periodic summarization. Although overwrite memory can in principle retain anything an append-only memory can, it must learn to carry each fact through every subsequent rewrite, which is difficult to learn from sparse outcome rewards; for interactive applications like web agents, we find that trained overwrite memories delete key information required by the trajectory, as well as corrective feedback received from the environment. We introduce AMBER (Append-only Memory Bank for Evidence Retention) - a simple and scalable framework where an agent jointly learns to reason, act, and write free-form memory, while an append-only rule guarantees retention by construction. This allows AMBER to be trained end-to-end with reinforcement learning from outcome rewards without the need for extensive curated SFT data. On WebArena Lite, AMBER improves average success over overwrite-based memory by 4.09 percentage points, increases the fraction of tasks solved in five repeated runs by 4.8 percentage points, and matches an overwrite baseline trained on substantially more expensive curated supervision. AMBER achieves these improvements while maintaining a practical token budget, providing a strong balance between context efficiency, task performance, and reliable long-horizon execution.
Comments29 pages, 11 figures