发表机构
Institute of Artificial Intelligence, Xiamen University; Key Laboratory of Digital Protection and Intelligent Processing of Intangible Cultural Heritage of Fujian and Taiwan (Xiamen University), Ministry of Culture and Tourism; Zhejiang University; Alibaba Group(厦门大学人工智能学院; 文化和旅游部闽台非物质文化遗产数字化保护与智能处理重点实验室(厦门大学); 浙江大学; 阿里巴巴集团)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对现有自演化记忆方法的信用分配问题,提出CHIME框架,通过维护规划与执行双记忆库并遵循先归因后记忆原则,在四个长视界智能体基准上优于多种基线,且记忆高效可迁移。
AI 中文摘要
规划是使智能体将复杂长视界任务分解为可管理步骤的核心能力。测试时搜索和基于训练的方法可提升规划能力,但会产生高推理成本或需要昂贵的训练数据。自演化记忆则将智能体交互结果中的可复用经验积累到外部记忆库中,使规划能力在推理时无需参数更新即可持续提升。然而,现有自演化记忆方法存在固有的信用分配问题:它们依赖最终任务结果作为反馈,但此类结果混淆了规划质量与执行错误、环境因素,因此积累的规划经验常存在偏差和噪声。为解决该问题,我们提出信用感知分层记忆演化(CHIME),这是一种自演化记忆框架,它维护独立的规划库与执行库,并遵循“先归因后记忆”原则:CHIME首先将每个任务结果归因于规划、执行、两者或都不,然后仅更新对应的记忆库。在四个长视界智能体基准上的大量实验表明,CHIME始终优于最先进的基于训练的方法和自演化记忆基线。进一步分析揭示了若干有趣发现:例如,CHIME用少得多的条目即可积累有效记忆;此外,学习到的记忆值忠实地反映了下游效用:高质量规划记忆比执行记忆更有价值;最后,积累的记忆可在骨干模型间有效迁移。代码将发布在此https URL。
英文摘要
Planning is a central capability that enables agents to decompose complex long-horizon tasks into manageable steps. Test-time search and training-based methods improve planning but incur high inference costs or require expensive training data. Self-evolving memory instead accumulates reusable experience from agent interaction outcomes into an external memory bank, so planning capability keeps improving at inference time without parameter updates. However, existing self-evolving memory methods share an inherent credit assignment problem: they rely on final task outcomes as feedback, but such outcomes conflate plan quality with execution errors and environmental factors, so the accumulated planning experience is often biased and noisy. To address this problem, we propose Credit-Aware Hierarchical Memory Evolution (CHIME), a self-evolving memory framework that maintains a separate planning bank and execution bank and follows an attribute-before-memorize principle: CHIME first attributes each task outcome to the plan, the execution, both, or neither, and then updates only the corresponding memory bank. Extensive experiments on four long-horizon agent benchmarks show that CHIME consistently outperforms state-of-the-art training-based and self-evolving memory baselines. Further analyses reveal several interesting findings. For example, CHIME accumulates effective memory with far fewer items. In addition, the learned memory values faithfully reflect downstream utility: high-quality planning memories are more valuable than execution memories. Finally, the accumulated memory effectively transfers across backbone models. Code will be released at https://github.com/ATH-MaaS/Marco-DeepResearch.
Comments7 figures, 7 tables