arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AECG:多智能体系统中的非对称经验整合与治理

AECG: Asymmetric Experience Consolidation and Governance In Multi-Agent Systems

Ao Tian, Jialong Liu, Daqi Zheng, Xin Sun, Mengting Li, Zhizhao Xiao, Zijian Huang, Honglei Wang, Zijian Hei, Yukun Yan

arXiv 2610.05176首次发表:更新:

发表机构

Beihang University; Wuhan University; ModelBest; Tsinghua University(北京航空航天大学; 武汉大学; ModelBest; 清华大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对多智能体系统中记忆污染和范围坍缩问题,提出AECG框架,通过非对称经验整合与治理,将记忆转化为动态可靠性治理循环,在多个基准上显著提升性能。

AI 中文摘要

基于大型语言模型(LLM)的多智能体系统日益依赖记忆将执行轨迹转化为可复用的程序性知识。然而,重复检索也使记忆错误持续存在:当过时、支持不足或虚假成功的程序成为未来推理的反复组成部分时,就会产生记忆污染。多智能体执行引入了额外的结构性风险。当程序性知识逃逸出其被证明有效的协调范围,并在不兼容的决策层级上被反复重用,使得局部错误影响下游决策的级联时,就会发生范围坍缩。同时,任务级失败提供的监督模糊,因为它们很少揭示哪个被召回的知识应负责。我们引入了AECG,一个用于多智能体系统的非对称经验整合与治理框架。AECG将记忆从静态经验存储转变为动态可靠性治理循环,保留协调范围,并使用多尺度、置信度感知的可靠性来检测退化。然后,它将退化与下游影响相结合,在有限的审查预算下优先处理高风险知识,应用有针对性的干预,并仅在配对重放后重新激活修订后的技能。在三个多智能体框架和四个基准测试中,AECG在12个框架-基准设置中的11个中取得了最佳分数,并比最强的竞争记忆方法提高了多达10.23个百分点;移除范围保留会使准确率降低多达16.89个百分点。因此,AECG将多智能体记忆从被动积累重新定义为可审计的可靠性治理。代码可在以下URL获取:https://this URL

英文摘要

Large language model (LLM)-based multi-agent systems increasingly rely on memory to transform execution trajectories into reusable procedural knowledge. Yet repeated retrieval also makes memory errors persistent: memory pollution arises when outdated, weakly supported, or spuriously successful procedures become recurring components of future reasoning. Multi-agent execution introduces an additional structural risk. Scope collapse occurs when procedural knowledge escapes the coordination scope in which it was shown effective and is repeatedly reused at incompatible decision levels, allowing local errors to influence cascades of downstream decisions. Meanwhile, task-level failures provide ambiguous supervision because they rarely reveal which recalled knowledge was responsible. We introduce AECG, a framework for asymmetric experience consolidation and governance for multi-agent systems. AECG turns memory from static experience storage into a dynamic reliability-governance loop, preserving coordination scope and using multi-scale, confidence-aware reliability to detect degradation. It then combines degradation with downstream impact to prioritize high-risk knowledge under a bounded review budget, applies targeted interventions, and reactivates revised skills only after paired replay. Across three multi-agent frameworks and four benchmarks, AECG achieves the best score in 11 of 12 framework--benchmark settings and improves over the strongest competing memory method by as much as 10.23 percentage points; removing scope preservation reduces accuracy by up to 16.89 points. AECG thereby reframes multi-agent memory from passive accumulation into auditable reliability governance. Code is available at https://github.com/fenhg297/AECG

Comments22 pages,5 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑