arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

学习保留什么:用于多智能体大语言模型系统高效协作的门控记忆路由

Learning What to Retain: Gated-Memory Routing for Efficient Collaboration in Multi-Agent LLM Systems

Rakibul Hasan Rajib, Mengxin Zheng, Qian Lou

arXiv 2609.00237首次发表:更新:

发表机构

University of Central Florida(中佛罗里达大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究针对多智能体LLM系统的协作协调问题,提出门控记忆路由方法,通过学习的门控机制优化记忆状态,在五组基准测试中准确率最优且推理成本降低,兼顾有效性与高效性。

AI 中文摘要

基于大语言模型(LLM)的多智能体系统通过协调多个智能体的配置与协作方式解决复杂推理问题,核心挑战在于让协调机制适配不断变化的协作状态。仅基于查询进行路由无法适配中间进展或错误,会损害准确性;基于完整执行历史进行路由可补充缺失的上下文,但会迫使后续决策处理所有先前步骤,包括冗余或低效用步骤,造成执行历史过载,推高成本。有效的协调需要能捕获有用进展且不积累冗余上下文的紧凑状态。我们提出门控记忆路由(Gated-Memory Routing),其每个决策以查询和学习到的执行记忆为条件:学习到的记忆写入门仅保留非冗余推理步骤,学习到的检索门为每个智能体提供紧凑、相关的子集,使每个决策都基于清晰、信息丰富的状态;系统每一步从该记忆中选择下一个角色和骨干,自适应停止控制器(Adaptive Halting Controller)在记忆包含足够回答证据时停止执行。在五个推理和代码生成基准测试中,我们的框架兼具有效性与高效性:平均准确率达到最优,比最强基线高出2.44个百分点,同时在HumanEval上的推理成本相对该基线降低31.9%。代码可在https URL获取。

英文摘要

Large language model (LLM)-based multi-agent systems tackle complex reasoning by orchestrating how multiple agents are configured and how they collaborate. A central challenge is to adapt orchestration to the evolving collaboration state. Routing from the query alone cannot adapt to intermediate progress or errors, which hurts accuracy. Routing from the complete execution history supplies this missing context, but forces later decisions to process every prior step, including redundant or low-utility ones. This creates an execution-history overload that inflates cost. Effective orchestration instead requires a compact state that captures useful progress without accumulating redundant context. We propose Gated-Memory Routing, which conditions each decision on the query and a learned execution memory. A learned Memory Write Gate commits only non-redundant reasoning steps, and a learned Retrieval Gate supplies each agent a compact, relevant subset, so every decision conditions on a clean, informative state. At each step, the system selects the next role and backbone from this memory, while an Adaptive Halting Controller stops execution once the memory contains sufficient evidence for answering. Across five reasoning and code-generation benchmarks, our framework is both effective and efficient: it attains the best average accuracy, exceeding the strongest baseline by 2.44 points, while reducing HumanEval inference cost by 31.9% relative to that baseline. Code is available at https://github.com/rajibrhasan/gated-memory-routing

CommentsAccepted to EMNLP 2026 (Main)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑