AI 中文总结
针对长 horizon LLM 智能体的内存-行动差距,提出感知功能的内存仲裁框架 MemArbiter,在 ALFWorld 上大幅提升行动成功率,减少失败行动相关问题。
AI 中文摘要
大型语言模型(LLM)智能体必须保留并利用跨步骤信息,才能在长 horizon 任务中连贯行动。现有方法提升了内存可访问性,但与行动相关的信息仍可能因形成、组织、优先级或呈现不佳,无法指导当前决策,我们将这种访问后失败称为内存-行动差距。我们提出 MemArbiter,一种感知功能的内存仲裁框架,用于解决该差距中由内存管理引发的部分。MemArbiter 将交互历史分解为原子项,将其组织为五个功能内存库,并结合库级需求、项级相关性、焦点-环境表示以及时间呈现门,动态控制内存显著性。我们在 ALFWorld 上以统一的每步内存预算,对比 Flat Retrieval 和 Flat Recency 评估 MemArbiter。采用开放权重的行动生成模型时,MemArbiter 在 500 和 750 token 预算下的成功率分别为 82.8% 和 92.5%,较最强基线分别提升 20.9 和 25.4 个百分点。它还改善了失败后恢复情况,减少了失败行动重复和状态-行动循环。这些结果表明,感知功能的内存仲裁能让可访问信息更有效地指导行动。
英文摘要
Large language model (LLM) agents must retain and use cross-step information to act coherently in long-horizon tasks. Existing methods improve memory accessibility, yet action-relevant information may still fail to guide the current decision because it is poorly formed, organized, prioritized, or presented. We call this post-access failure the Memory-Action Gap. We propose MemArbiter, a function-aware memory arbitration framework that addresses the memory-management-induced component of this gap. MemArbiter decomposes interaction histories into atomic items, organizes them into five functional Memory Banks, and combines bank-level demand, item-level relevance, focal-ambient representations, and a temporal presentation gate to dynamically control memory salience. We evaluate MemArbiter on ALFWorld against Flat Retrieval and Flat Recency under unified per-step memory budgets. With an open-weight action-generation model, MemArbiter achieves success rates of 82.8% and 92.5% under 500- and 750-token budgets, outperforming the strongest baseline by 20.9 and 25.4 percentage points, respectively. It also improves post-failure recovery and reduces failed-action repetition and state-action recurrence. These results show that function-aware memory arbitration enables accessible information to guide actions more effectively.
Comments9 pages, 3 figures, 5 tables