arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

多样性指标真的在衡量多样性吗?对大语言模型集成中多数投票增益的能力控制审计

Are Diversity Metrics Measuring Diversity? A Capability-Controlled Audit of Majority-Vote Gain in LLM Ensembles

Donghwan Kim

arXiv 2607.20768首次发表:更新:

发表机构

Aidentyx Inc.(艾登泰克斯公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

探讨大语言模型集成中多样性指标与多数投票增益的关系,在能力控制下审计五个多样性指标,发现潜在互补性普遍、指标与准确率共线且统计量不可分离,揭示了多样性指标在衡量多样性方面的问题及相关关联。

AI 中文摘要

人们普遍认为,大语言模型的多数投票得益于多样性,并且使用多样性指标来选择要组合的模型。我们探讨了五个这样的指标是在追踪多样性还是主要重新表达能力,并在明确的能力控制下,将它们作为30个大语言模型在MMLU-Pro(29个在TruthfulQA)上的31900个子集上相对于最佳成员的多数投票增益的预测指标进行审计。有三个发现。首先,潜在互补性普遍存在:神谕增益在100%的子集中为正,但简单投票仅在9.98%的所有标准大小为3的子集中击败最强成员(在保留最佳选择的情况下为18.71%);合并的大小为2至4的比率为1.27%,部分反映了确定性的偶数大小投票行为。其次,表示联合正确性的代理指标(严格多样性)与1减去平均准确率几乎共线(大小为3时斯皮尔曼相关系数为+0.991 / +0.988);原始的多样性增益关联与能力紧密纠缠,并且除了一个例外,在控制下不稳定。第三,三个线性列联表统计量在代数上不可分离;在进行能力控制后,经验上稳定的剩余部分是一个适度的残余成对共同失败关联,其中更多的共享错误对应于更低的增益。这个方向是稳健的,但其大小取决于配置。将严格多样性、分歧和双故障视为独立预测指标的联合原始空间线性回归在构造上是秩亏的。

英文摘要

Majority voting over LLMs is widely assumed to benefit from diversity, and diversity measures are used to choose which models to combine. We ask whether five such measures track diversity or mainly re-express capability, auditing them as predictors of majority-vote gain over the best member across 31,900 subsets of 30 LLMs on MMLU-Pro (29 on TruthfulQA) under explicit capability controls. Three findings emerge. First, latent complementarity is ubiquitous: oracle gain is positive in 100% of subsets, yet simple voting beats the strongest member in only 9.98% of all canonical size-3 subsets (18.71% with held-out best selection); the pooled size-2-4 rate is 1.27%, partly reflecting deterministic even-size voting behavior. Second, a joint-correctness proxy (strict diversity) is nearly collinear with one minus mean accuracy (size-3 Spearman rho = +0.991 / +0.988); raw diversity-gain associations are strongly capability-entangled and, with one exception, unstable under control. Third, three linear contingency-table statistics are algebraically non-separable; after capability control, the empirically stable remainder is a modest residual pairwise co-failure association in which more shared error corresponds to lower gain. This direction is robust, but its magnitude is configuration-dependent. Joint rawspace linear regressions treating strict diversity, disagreement, and double-fault as independent predictors are rank-deficient by construction.

Comments10 figures, 9 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑