arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

候选供给与答案选择塑造多智能体系统中大语言模型(LLM)评判的价值

Candidate supply and answer selection shape the value of LLM judging in multi-agent systems

Jia-Hao Ji, Sijie Li, Jiabei Cheng, Zixi She, Jin-Tai Yu, Zhiyuan Yuan

arXiv 2608.25937首次发表:更新:

AI 中文总结

该研究针对多智能体系统中LLM评判的价值,通过分析多组基准数据集,发现结合答案频率与LLM评判可提升答案准确率,为设计多智能体架构提供了诊断基础。

AI 中文摘要

多智能体系统(MAS)有时已具备正确作答的潜力,却仍会报告错误答案。由于生成、通信及最终答案选择规则通常会同时变化,解释这一结果十分困难。我们将多智能体推理概念化为一个包含候选生成、同伴通信与终端选择的进化流程,其中缺乏质量控制的共识会表现出模因漂移模式。我们研究两个问题:(1)当大语言模型(LLM)评判为多智能体系统生成的候选提供答案正确性信号,从而产生有效选择压力的情况;(2)使用该信号提升报告答案的情况。为绘制评判可靠性图谱,我们分析了来自MMLU-Pro、GPQA、MedXpertQA和MuSR的15336个问题,对Humanity's Last Exam则单独分析。为测试这些规则,我们重放了从五个基准的16278个问题中抽取的81390个固定候选池。我们报告三项发现:(1)正确答案通常已存在于生成的候选中,但系统仍可能收敛并报告错误答案;(2)评判可靠性并非模型的固定属性,而是随任务、生成器及正确答案的稀有程度变化;(3)将答案频率与评判评估结合,仅改变了最终答案选择规则,使准确率从63.82%提升至70.82%-70.95%,主要是拯救了被流行错误数量超过的正确答案。在所研究的系统中,生成更多候选的价值取决于这些额外样本是否使正确答案存在、频繁或可识别。通过隔离生成、识别与选择,这些发现为设计多智能体架构建立了诊断基础,以保护生成的正确答案不丢失。

英文摘要

Multi-agent systems (MAS) sometimes already have the potential to answer correctly, but still report a wrong answer. Explaining this outcome is difficult because generation, communication and final answer-selection rules usually change simultaneously. We conceptualize multi-agent reasoning as an evolutionary pipeline of candidate generation, peer communication and terminal selection, wherein consensus without quality control can exhibit patterns of memetic drift. We study two questions: (1) when an LLM judge provides effective selection pressure by supplying a signal of answer correctness for candidates generated in a multi-agent system, and (2) when using that signal improves the reported answer. To map judge reliability, we analysed 15,336 questions from MMLU-Pro, GPQA, MedXpertQA and MuSR, with Humanity's Last Exam analysed separately. To test these rules, we replayed 81,390 fixed candidate pools drawn from 16,278 questions across five benchmarks. We report three findings. (1) A correct answer is often already present among the generated candidates, but the system can still converge on and report a wrong answer. (2) Judge reliability is not a fixed trait of the model, but varies with the task, the generator and how rare the correct answer is. (3) Combining answer frequency with the judge's evaluation changed only the final answer-selection rule and raised accuracy from 63.82% to 70.82-70.95%, primarily by rescuing correct answers that were outnumbered by popular errors. In the systems studied here, the value of generating more candidates depends on whether those extra samples make correct answers present, frequent or recognisable. By isolating generation, recognition and selection, these findings establish a diagnostic basis for designing multi-agent architectures that protect generated correct answers from being lost.

Comments11 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑