发表机构
The University of Melbourne; Sri Lanka Institute of Information Technology(墨尔本大学; 斯里兰卡信息技术学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对多智能体意见冲突问题,提出基于贝叶斯反向推理构建反向后验作为无标签锚点,利用Jensen-Shannon散度评估跨路径一致性,实现硬选择、软重加权和对数线性融合三种策略,在DDXPlus数据集上显著提升集体决策性能。
AI 中文摘要
当多个大语言模型(LLM)智能体产生相互矛盾的答案时,决策过程决定了智能体的多样性是提升性能还是仅仅加剧共享错误。现有的集体决策方法,包括投票、选举规则和LLM裁判,都依赖于正向推理:它们将证据单向映射到标签。尽管这些方法能够结合多样化的正向轨迹,但它们仍然聚合了共享这种证据到标签分解的估计,并可能在正向池内继承相关错误。因此,我们通过从显式似然出发的贝叶斯反向推理,为每个实例构建一个反向后验。正向和反向后验提供了对底层后验的不同分解近似。由于来自不同分解的估计可能更少倾向于共享相同的错误,我们使用Jensen-Shannon散度根据跨路径一致性对智能体进行排序。这种跨路径一致性信号支撑了三种策略:硬选择(MinJS)、软重新加权(FwdJS)和对数线性融合(LogLin)。在DDXPlus数据集上跨五个LLM骨干网络进行评估,我们提出的策略显示出持续改进:MinJS在所有骨干网络上均优于随机选择,FwdJS通常优于最强基线,而LogLin在评估方法中取得了最佳性能,其在智能体意见不一致的子集上增益最大。尽管其独立准确性较弱,反向后验作为比仅正向替代方案更有用的锚点,为集体决策提供了互补信息。当有标签数据可用时,一个轻量级的两阶段校准可以进一步细化反向锚点并提升聚合性能。
英文摘要
When multiple LLM agents yield conflicting answers, the decision-making process dictates whether agent diversity improves performance or merely compounds shared errors. Existing collective decision-making methods, including voting, electoral rules, and LLM judges, rely on forward reasoning: they map evidence to labels in one direction. Although these methods can combine diverse forward traces, they still aggregate estimates that share this evidence-to-label factorization and can inherit correlated errors within the forward pool. We therefore construct a reverse posterior for each instance through Bayesian backward reasoning from an explicit likelihood. The forward and reverse posteriors provide differently factorized approximations of the underlying posterior. Because estimates from different factorizations may tend to share the same error less often, we use Jensen-Shannon divergence to rank agents by cross-path consistency. This cross-path consistency signal underlies three strategies: hard selection (MinJS), soft reweighting (FwdJS), and log-linear fusion (LogLin). Evaluated on DDXPlus across five LLM backbones, our proposed strategies show consistent improvements: MinJS outperforms random selection across all backbones, FwdJS generally improves over the strongest baseline, and LogLin achieves the best performance among the evaluated methods, with its largest gains on the subset where the agents disagree. Despite its weaker standalone accuracy, the reverse posterior serves as a more useful anchor than forward-only alternatives, providing complementary information for collective decision-making. When labeled data are available, a lightweight two-stage calibration can further refine the reverse anchor and improve aggregation performance.