arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

预言机差距与信号保真度:一种用于测试时协作的固定池诊断方法

Oracle Gap and Signal Fidelity: A Fixed-Pool Diagnostic for Test-Time Collaboration

Jie Hu

arXiv 2607.17531首次发表:更新:

发表机构

Research Institute of China Telecom(中国电信研究院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究测试时协作收益不均衡问题,将固定候选池选择器或验证器净收益分解为可测量因素,通过实验表明收益受预言机差距和信号保真度限制,框架可在协作前进行预部署诊断,估计相关指标。

AI 中文摘要

测试时协作,包括自一致性、N选最佳、批评模型和验证管道,通常被认为能广泛改善大语言模型推理,但其收益不均衡且有时为负。我们探讨何时应期望无训练协作有所帮助。对于固定候选池,我们将选择器或验证器的净收益分解为可测量因素:可恢复质量、验证信号覆盖、条件选择质量以及对已正确输出的损害。这将协作重新构建为候选选择问题,而非多智能体拓扑的固有属性。在LiveCodeBench、MATH Level - 5难题和GPQA - Diamond上,收益首先受预言机差距限制,然后受信号保真度限制,我们将其直接测量为验证器裁决与官方标签之间的候选级一致性。在LiveCodeBench上,公共测试验证器(MCC 0.825)比首次采样基线提高了8.14个百分点;生成测试验证器(MCC 0.248)提高了2.70个百分点,与大语言模型选择器无统计学差异,但对已正确输出的损害率接近零,而选择器的损害率为4.69%。在MATH上,符号答案等效选择器比自一致性高出4.67个百分点,而大语言模型选择器则为负。在GPQA - Diamond上,可恢复质量仅为3.03%,87.54%的候选池答案相同;较弱模型的池进一步缩小,表明预言机差距是任务、模型和采样配置的联合属性。我们的框架产生了一种实用的预部署诊断方法:在投资协作之前,估计预言机差距,然后测量覆盖、信号保真度和损害。

英文摘要

Test-time collaboration, including self-consistency, best-of-N selection, critic models, and verifier pipelines, is often credited with broadly improving LLM reasoning, yet its gains are uneven and sometimes negative. We ask when training-free collaboration should be expected to help. For a fixed candidate pool, we decompose a selector or verifier's net gain into measurable factors: recoverable mass, verification-signal coverage, conditional selection quality, and harm to already-correct outputs. This reframes collaboration as a candidate-selection problem rather than as an intrinsic property of a multi-agent topology. Across LiveCodeBench, MATH Level-5 hard subjects, and GPQA-Diamond, gains are bounded first by the oracle gap and then by signal fidelity, which we measure directly as candidate-level agreement between verifier verdicts and official labels. On LiveCodeBench, a public-test verifier (MCC 0.825) gains +8.14 percentage points (pp) over a first-sample baseline; a generated-test verifier (MCC 0.248) improves by +2.70pp and is not statistically distinguishable from an LLM selector, but operates at near-zero harm versus the selector's 4.69% harm rate. On MATH, a symbolic answer-equivalence selector beats self-consistency by +4.67pp, while LLM selectors are negative. On GPQA-Diamond, recoverable mass is only 3.03% and 87.54% of candidate pools are answer-identical; a weaker model's pools shrink both further, suggesting that oracle gap is a joint property of task, model, and sampling configuration. Our framework yields a practical pre-deployment diagnostic: estimate the oracle gap, then measure coverage, signal fidelity, and harm before investing in collaboration.

Comments13 pages, 2 figures, 8 tables. Code: https://github.com/AmGarfield/OracleGap

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑