辩论何时有帮助:多智能体推理中的提案供给与验证感知读出
When Debate Helps: Proposal Supply and Verification-Aware Readout in Multi-Agent Reasoning
浏览论文内容
中文总结 AI 辅助
本研究提出可恢复余量与潜在验证辩论模型,结合覆盖选择的神经丛智能体社会,证明提案覆盖与真值敏感证据使用是辩论超越多数投票的互补条件。
中文摘要 AI 辅助
多智能体辩论可以改善推理,但往往无法超越简单多数投票。我们认为,成功的辩论需要两个不同的机制:提案供给必须呈现出一个正确答案,而读出必须在投票遗漏该答案时识别出它。我们通过可恢复余量来形式化第一个要求,该余量衡量了正确答案可用但多数答案错误的情况。对于第二个要求,我们开发了潜在验证辩论(LVD),这是一种核算模型,其中候选提案在最终生成前接收特定于答案的验证证据。受控的固定提案干预以等效同行支持单位估计这种潜在效应,并表明正确的证据会改变答案概率和生成的决策,而提案供给保持不变。为了改善提案供给,我们使用标记和无标记覆盖目标从神经丛智能体构建社会。在两个主干和匹配预算的推理基准上,覆盖选择的社增加了互补的提案供给,并在重复随机评估中提高了总体准确率。轮级控制进一步表明,交互带来的收益超过了直接将相同终结器应用于初始提案。这些结果确定了提案覆盖和真值敏感的证据使用是辩论优于投票的互补条件。代码可在该https URL获取。
英文摘要
Multi-agent debate can improve reasoning, yet often fails to beat simple majority voting. We argue that successful debate requires two distinct mechanisms: proposal supply must surface a correct answer, and readout must identify that answer when voting misses it. We formalize the first requirement through recoverable headroom, which measures cases where a correct proposal is available but the majority answer is wrong. For the second, we develop Latent Verification Debate (LVD), an accounting model in which candidate proposals receive answer-specific verification evidence before final generation. Controlled fixed-proposal interventions estimate this latent effect in equivalent peer-support units and show that correct evidence changes answer probabilities and generated decisions while proposal supply remains fixed. To improve proposal supply, we construct societies from neural-thicket agents using labeled and label-free coverage objectives. Across two backbones and matched-budget reasoning benchmarks, coverage-selected societies increase complementary proposal supply and improve aggregate accuracy in repeated stochastic evaluations. Round-level controls further show that interaction provides gains beyond applying the same finalizer directly to the initial proposals. These results identify proposal coverage and truth-sensitive evidence use as complementary conditions for debate to outperform voting. Code is available at https://github.com/Wang-ML-Lab/when-debate-helps.
发表机构
- University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)
- Rutgers University(罗格斯大学)
机构由 AI 辅助整理,请以论文原文为准。