arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

量化思维多样性:加权大语言模型集成提升的预测定律

Quantifying Diversity of Thought: A Predictive Law of Weighted LLM Ensemble Lift

Junade Ali

arXiv 2607.17384首次发表:更新:

发表机构

The Alan Turing Institute(艾伦·图灵研究所)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究量化思维多样性在大语言模型集成中提升效果的预测定律,通过原理推导得出计算提升的启发式方法及相关指标,经多基准测试验证,该方法能有效预测集成性能,且准确性调整后的指标表现更优。

AI 中文摘要

本文提供了一个经实验验证的形式定律,用于计算思维多样性在大语言模型(LLM)集成中所带来的提升。从基本原理出发,我们将LLM集成提升精确分解为救援质量和损害质量,从而得出一个计算提升的简洁启发式方法。从中提取出预测集成性能的指标:一个准确性调整后的正确性相关性$\phi_{\mathrm{adj}}$,以及该对的准确性差距和集体准确性。我们在两个研究生水平的科学基准测试以及一个新颖的智能网络安全基准测试上,对十个开放权重模型的767,520次推理进行了测试。该启发式方法在SuperGPQA上以40:60的投票分割进行一次校准后,在校准集上预测提升的斯皮尔曼$\rho$为0.84,并且在系数冻结的情况下,转移到校准中从未使用过的两个数据集上(在GPQA Diamond上$\rho$为0.51,在法医任务上为0.84),同时测量的交换质量在整个过程中与实际提升的$R^2\geq0.96$。原始的$\phi$几乎没有预测能力(整个过程中$R^2\leq0.09$);准确性调整后的$\phi_{\mathrm{adj}}$明显更优(在SuperGPQA上$R^2 = 0.67$),并且结合这些指标的启发式方法是三个数据集上最稳定的预池预测器。

英文摘要

This paper provides an experimentally verified formal law for calculating the uplift that diversity of thought provides in Large Language Model (LLM) ensembles. From first principles, we derive an exact decomposition of LLM ensemble lift into rescue and damage masses, which yields a compact heuristic for calculating uplift. From this we extract the metrics which predict ensemble performance: an accuracy-adjusted correctness correlation, $ϕ_{\mathrm{adj}}$, together with the accuracy gap and collective accuracy of the pair. We test the law on 767,520 inferences from ten open-weight models over two graduate-level science benchmarks, together with a novel agentic cybersecurity benchmark in which each model conducts digital-forensics investigations by multi-turn tool use in a network-isolated sandbox (23,520 graded trials including abstentions); all votes are released openly. Calibrated once on SuperGPQA at a 40:60 vote split, the heuristic predicts lift on the calibration set with Spearman's $ρ=0.84$ and, with its coefficients frozen, transfers to two datasets never used in calibration ($ρ=0.51$ on GPQA Diamond and $0.84$ on the forensic tasks), whilst the measured swap mass tracks realised lift with $R^2\ge 0.96$ throughout. Raw $ϕ$ has almost no predictive power ($R^2\le 0.09$ throughout); the accuracy-adjusted $ϕ_{\mathrm{adj}}$ is markedly superior ($R^2=0.67$ on SuperGPQA), and the heuristic combining these metrics is the most stable pre-pooling predictor across the three datasets.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑