arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

同质面板LLM辩论中的坍缩与修正测量

Measuring Collapse and Correction in Homogeneous-Panel LLM Debate

Xin Li, Mengbing Liu, Chau Yuen

arXiv 2609.35279首次发表:更新:

发表机构

Nanyang Technological University(南洋理工大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对同质面板LLM辩论,提出可审计协议区分坍缩与修正,通过6,925场辩论揭示防坍缩策略可能损失修正,并发布重放工具以支持公平比较。

AI 中文摘要

多智能体大语言模型(LLM)辩论通常通过最终答案是否改善来评估,但变化并不一定意味着改善:同一场讨论既可能拯救最初错误的大多数,也可能摧毁最初正确的大多数。标准的最终准确率评估混淆了这两种相反的机制。我们提出了一种针对多项选择题(MCQs)的同质辩论的可审计协议,将每次运行记录为包含坍缩、修正、起始和带符号干预效用的转换账本。在6,925场MMLU-Pro辩论中,该协议识别出253次坍缩,并生成了一个并行的修正账本,改变了干预措施应被评判的方式。重放实验揭示了核心权衡:在等权重下,一种留一模型出的探针门控冻结可防止29次坍缩,但会损失108次修正,因此仅防止坍缩可能推荐错误的策略。一个紧凑的辩论前8探针筛选是一个分诊信号:其与条件坍缩风险的未调整族级关联很高(G=7,Spearman rho=0.893,精确双侧p=0.0123),但初始多数准确率是接近的对照(rho=0.821;族偏rho=0.767,p=0.0877),因此我们不将其视为校准或能力调整的预测。轮次级追踪将许多坍缩定位到第一轮辩论,其中早期分歧可能先于有害级联和有用的恢复。我们发布了可重放的模式、编码器、审计、成本卡和零API重建脚本,以便未来的模型-脚手架行可以在相同的分母和带符号效用账本下进行比较。

英文摘要

Multi-agent large language model (LLM) debate is often evaluated by whether final answers improve, but movement is not necessarily improvement: the same discussion can rescue an initially wrong majority or destroy an initially correct one. Standard final-accuracy evaluations conflate these opposing mechanisms. We introduce an auditable protocol for homogeneous debate on multiple-choice questions (MCQs) that records each run as a transition ledger over collapse, correction, onset, and signed intervention utility. On 6,925 MMLU-Pro debates, the protocol identifies 253 collapses and a parallel correction ledger that changes how interventions should be judged. Replay experiments reveal the central tradeoff: a leave-one-model-out probe-gated freeze prevents 29 collapses but loses 108 corrections under equal weights, so collapse prevention alone can recommend the wrong policy. A compact pre-debate 8-probe screen is a triage signal: its unadjusted family-level association with conditional-collapse risk is high (G=7, Spearman rho=0.893, exact two-sided p=0.0123), but initial-majority accuracy is a close comparator (rho=0.821; family partial rho=0.767, p=0.0877), so we do not treat it as calibrated or capability-adjusted prediction. Round-level traces localize many collapses to the first debate round, where early disagreement can precede both harmful cascades and useful recovery. We release replayable schemas, coders, audits, cost cards, and zero-API rebuild scripts so future model-scaffold rows can be compared under the same denominators and signed utility ledger.

CommentsAccepted at NeurIPS 2026 (Evaluations and Datasets Track). Project page: https://lixin.ai/DebateLedger. Code: https://github.com/LiXin97/DebateLedger

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑