发表机构
University of Milano–Bicocca(米兰比可卡大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对多LLM委员会在持续对抗者下的置信度校准与稳健性问题,提出贝叶斯辩证论证(BDA),通过建模类型化动作和每智能体可靠性,实现校准后验概率并反转不可靠智能体,在零成本方法中校准最优且鲁棒。
AI 中文摘要
多LLM委员会让多个大型语言模型(LLM)就一个问题进行审议,并返回答案及置信度估计。随着这些系统越来越多地用于推理,该置信度应代表校准的“正确概率”,且当某些智能体持续不可靠时,决策应保持稳健。现有的委员会聚合方法在这两方面均失败:其置信度估计衡量的是决断性而非正确性,且无法识别或折扣持续不可靠的智能体。我们引入贝叶斯辩证论证(BDA),将委员会的“类型化”动作——谁提出、质疑或承认哪个答案——视为经典标注者模型的观测,该模型具有“每智能体”可靠性。这一表述将多智能体审议重新定义为可靠性估计问题,利用审议轨迹在持续对抗行为下推断智能体可靠性。通过根据推断的智能体可靠性对证据加权,BDA在候选答案上产生校准的后验概率,同时允许持续不可靠的智能体被反转而非仅仅被否决。在二元和多类基准上,BDA在零成本委员会聚合方法中实现了最佳校准,无需额外LLM调用,并在持续对抗联盟下提高了稳健性,同时在干净环境中保持竞争力。
英文摘要
A multi-LLM \emph{council} lets several large language models (LLMs) deliberate on a question and return an answer together with a confidence estimate. As these systems become increasingly used for reasoning, that confidence should represent a calibrated \emph{probability of being correct}, and the decision should remain robust when some agents are persistently unreliable. Existing \emph{council aggregation} methods fail on both fronts: their confidence estimates measure decisiveness rather than correctness, and they cannot identify or discount persistently unreliable agents. We introduce Bayesian Dialectical Argumentation (BDA), which treats the council's \emph{typed} moves---who proposed, challenged, or conceded which answer---as observations of a classical annotator model with \emph{per-agent} reliabilities. This formulation recasts multi-agent deliberation as a reliability estimation problem, using the deliberation trace to infer agent reliability under persistent adversarial behavior. By weighting evidence according to inferred agent reliability, BDA yields calibrated posterior probabilities over candidate answers while allowing persistently unreliable agents to be inverted rather than merely outvoted. Across binary and multi-class benchmarks, BDA achieves the best calibration among zero-cost council aggregation methods, requiring no additional LLM calls, and improves robustness under persistent adversarial coalitions while remaining competitive in clean settings.