元审核员:用元认知赋能多智能体辩论
Meta-Moderator: Empowering Multi-Agent Debate with Meta-Cognition
浏览论文内容
中文总结 AI 辅助
该研究提出Meta-Moderator可学习框架,将多智能体辩论的审核视为元认知过程,经结果驱动策略优化训练,在五个基准测试中性能优于现有决策层,能动态调控辩论并减少错误聚合。
中文摘要 AI 辅助
多智能体辩论可通过引出多样假设与批判来提升大语言模型的推理能力,但其性能常受限于薄弱的审核机制。现有常见流程依赖固定预算、基于一致性的停止条件或未经过训练的评判者,导致冗余的审议过程与不可靠的证据聚合。我们将审核视为一种元认知过程,其功能为监控辩论效用、控制审议流程并裁决最终答案;同时引入Meta-Moderator,这是一种可学习的框架,能动态调控辩论并决定何时确定最终答案。Meta-Moderator通过结果驱动的策略优化独立于辩论者进行训练,使辩论调控成为一种显式能力,而非提示的附带效果。在五个基准测试中,Meta-Moderator的性能优于广泛使用的决策层,且可跨任务与系统配置迁移。进一步分析显示,它能更有选择性地分配辩论资源,并在出现有信息量的假设后减少错误聚合。
英文摘要
Multi-agent debate can improve large language model reasoning by eliciting diverse hypotheses and critiques, yet its performance is often constrained by weak moderation. Common pipelines rely on fixed budgets, agreement-based stopping, or untrained judges, leading to redundant deliberation and unreliable evidence aggregation. We cast moderation as a meta-cognitive process, monitoring debate utility, controlling deliberation, and adjudicating a final answer, and introduce Meta-Moderator, a learnable framework that dynamically regulates debate and decides when to finalize an answer. Meta-Moderator is trained independently of the debaters via outcome-driven policy optimization, making debate regulation an explicit capability rather than an incidental effect of prompting. Across five benchmarks, Meta-Moderator outperforms widely used decision layers and transfers across tasks and system configurations. Further analyses show that it allocates debate more selectively and reduces mis-aggregation after informative hypotheses appear.
发表机构
- The Hong Kong Polytechnic University(香港理工大学)
- Sichuan University(四川大学)
机构由 AI 辅助整理,请以论文原文为准。