先一致后多样:异构语言模型协作的优先验证互补机制
Agreement Before Diversity: Verification-First Complementarity for Heterogeneous Language-Model Coordination
浏览论文内容
中文总结 AI 辅助
本文提出异构语言模型协作的Agreement-Before-Diversity优先验证机制,通过解耦候选余量与替换权限,在LiveCodeBench和GPQA-Diamond数据集上实现了优于基线的准确率,同时为门控机制证明了两个精确恒等式。
中文摘要 AI 辅助
异构语言模型集成拓展了候选响应空间,但缺乏明确的准则判定新生成答案是否应替代已支持的答案。本文将候选余量与替换权限解耦,使替换权限成为显式、可审计的对象,提出的方法为Agreement-Before-Diversity(ABD),是一种冻结的无标签决策规则:若两个额外可信样本在固定等价关系下证实锚定答案,则保留该锚定答案;否则用异构合成结果替换。针对该门控机制,本文证明两个精确恒等式:第一个表明,相对于无条件合成的准确率差距由一致覆盖度和锚定答案在受保护子集上的优势共同决定;第二个表明,相对于从不合成的差距反映了授权恢复与授权破坏的对比。两个恒等式均不假设独立性或校准置信度,预期推理成本约为8减去调用次数的五倍。在盲测精确ID评估中,ABD在完整LiveCodeBench-v6数据集上达到59.43%(对比Single9的52.57%和HAC的52.00%,样本量n=175),在未触及的GPQA-Diamond子集上达到75.00%(两个对照组均为72.78%,样本量n=180)。此外,这些恒等式将所有聚合差异定位到可枚举的受保护层:LiveCodeBench的3个受保护案例中无不一致项,覆盖度先验限定门控贡献为1.71个百分点;GPQA-Diamond的132个案例中有13个不一致项,对照组为8个;冻结锚定扰动下的71个案例中有12个不一致项,对照组为0个。多样性提供潜力,验证结构提供权限。
英文摘要
Heterogeneous language-model ensembles expand the space of candidate responses, yet they lack a principled criterion for when a newly generated answer should supersede an already supported one. We decouple candidate headroom from replacement authority, rendering the latter as an explicit, auditable object. Our proposed method, Agreement-Before-Diversity (ABD), is a frozen, label-free decision rule: an anchor answer is retained if two additional trusted samples corroborate it under a fixed equivalence relation; otherwise, it is replaced by a heterogeneous synthesis. For this gating mechanism, we prove two exact identities. The first shows that the accuracy gap relative to unconditional synthesis is determined jointly by the agreement coverage and the anchor's advantage on the protected subset. The second shows that the gap relative to never synthesizing reflects a contrast between authorized recovery and authorized destruction. Neither identity assumes independence or calibrated confidence, and the expected inference cost is approximately eight minus five times the coverage in number of calls. Under blind, exact-ID evaluation, ABD achieves 59.43% on the complete LiveCodeBench-v6 (vs. 52.57% for Single9 and 52.00% for HAC; n = 175) and 75.00% on an untouched GPQA-Diamond split (both controls at 72.78%; n = 180). Furthermore, these identities localize every aggregate difference to an enumerable protected stratum: no discordant items occur among the 3 protected cases on LiveCodeBench, where coverage bounds the gate's contribution to 1.71 points a priori; 13 versus 8 discordant cases among 132 on GPQA-Diamond; and 12 versus 0 among 71 under a frozen anchor perturbation. Diversity supplies potential; verification structure supplies authority.