记忆与重加权:利用经验记忆与置信度估计增强多智能体辩论
Remember and Reweight: Enhancing Multi-Agent Debate with Experience Memory and Confidence Estimation
浏览论文内容
中文总结 AI 辅助
该研究针对多智能体辩论的共享误解问题,提出R²-MAD框架,通过经验记忆与置信度估计校准概念先验、调节同伴影响,在多个基准测试上实现性能提升。
中文摘要 AI 辅助
多智能体辩论(MAD)通过让多个智能体通过讨论迭代优化其回答,提升了大语言模型的推理能力。然而,MAD存在一个关键漏洞,即共享误解:当多数智能体最初收敛于错误答案时,辩论过程往往会放大而非纠正该错误。现有方法主要解决同伴偏差问题,但未处理智能体固有的有偏概念先验。为缓解这一系统性缺陷,我们提出R²-MAD(多智能体辩论的记忆与重加权框架),该框架为智能体配备从过往辩论中积累的经验记忆。R²-MAD通过两种互补机制干预两类失败模式:一种辩论状态感知检索策略,基于当前共识水平检索相关历史证据,动态校准概念先验;随后,这些检索到的经验为估计各智能体的可靠性提供依据,进而生成置信度权重以调节同伴影响。在多个基准测试上的实验表明,R²-MAD相较于现有单智能体及MAD基线实现了持续提升。
英文摘要
Multi-agent debate (MAD) improves the reasoning capabilities of large language models by having multiple agents iteratively refine their responses through discussion. However, MAD suffers from a critical vulnerability known as shared misconception: when a majority of agents initially converge on an incorrect answer, the debate process tends to amplify rather than correct the error. Existing methods primarily address peer skew but leave the agents' inherently biased concept priors unaddressed. To mitigate this systematic weakness, we propose R$^2$-MAD (Remember and Reweight for Multi-Agent Debate), a framework that equips agents with an experience memory accumulated from past debates. R$^2$-MAD intervenes on both failure modes through two complementary mechanisms: A debate-state-aware retrieval policy dynamically calibrates the concept prior by retrieving relevant historical evidence based on the current consensus level. Then these retrieved experiences provide a basis for estimating per-agent reliability, yielding confidence weights to modulate peer influence. Experiments on various benchmarks show that R$^2$-MAD achieves consistent improvements over existing single-agent and MAD baselines.
发表机构
- Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所)
- University of Chinese Academy of Sciences(中国科学院大学)
- University College London(伦敦大学学院)
机构由 AI 辅助整理,请以论文原文为准。