arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

MACE:面向多智能体系统的记忆-智能体协同进化与自适应记忆图

MACE: Memory-Agent Co-Evolution with Adaptive Memory Graphs for Multi-Agent Systems

Kairui Yang, Minghao An, Xunkai Li, Ziheng Yi, Zekai Chen, Guangyuan He, Rong-Hua Li

arXiv 2609.21533首次发表:更新:

AI 中文总结

MACE提出记忆-智能体协同进化框架,通过执行反馈自适应组织记忆并优化智能体使用,在八个基准上以81.11%平均分超越最强基线SAGE。

AI 中文摘要

基于大语言模型的多智能体系统生成协作轨迹,记录智能体如何规划任务、验证中间结果以及修复失败。重用这些流程需要保留动作的前提条件以及后续智能体所需的输出。我们的实证研究表明,将这些依赖关系分组为功能性记忆单元可提高其保留率,而连接单元则能提高任务所需的单元和链接的联合检索率。即使在每种组合的内容跨格式固定不变的情况下,指令与清单之间偏好的单元组合也会发生变化。根据每种组合与格式配对的结果更新选择,优于分别对组合和格式进行评分。这些发现催生了MACE,一个通过执行反馈自适应调整记忆组织和智能体记忆使用的记忆-智能体协同进化框架。其MemGoG结构将功能单元表示为相关条件、动作和输出的子图,并通过支持、冲突和修复关系连接它们。MACE Loop在记忆预算内选择任务相关单元和关系,并为每个智能体提供当前操作的指令或清单。它记录所选单元、呈现格式、智能体输出和任务结果,以更新单元得分和关系用于检索,并为后续呈现选择提供信息。在八个基准测试中,MACE以平均81.11%的得分优于十个基线,而最强基线SAGE的得分为78.97%。

英文摘要

LLM-based multi-agent systems generate collaboration traces that record how agents plan tasks, verify intermediate results, and repair failures. Reusing these procedures requires preserving an action's prerequisites and the outputs needed by subsequent agents. Our empirical studies show that grouping these dependencies into functional memory units improves their retention, while connecting units increases retrieval of the units and links jointly required by a task. The preferred combination of units also changes between instructions and checklists, even when each combination's content is fixed across formats. Updating choices from the outcomes of each combination and format pairing outperforms scoring combinations and formats separately. These findings motivate MACE, a memory-agent co-evolution framework that adapts memory organization and agent memory use through execution feedback. Its MemGoG structure represents functional units as subgraphs of related conditions, actions, and outputs, connecting them through support, conflict, and repair relations. MACE Loop selects task-relevant units and relations within a memory budget and provides each agent with instructions or checklists for its current operation. It records the selected units, presentation formats, agent outputs, and task outcomes to update unit scores and relations for retrieval and inform subsequent presentation choices. Across eight benchmarks, MACE outperforms ten baselines with an average score of 81.11%, compared with 78.97% for the strongest baseline, SAGE.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑