发表机构
Stony Brook University(石溪大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究针对LLM投票中的发现-决策差距,提出可恢复性阈值与精确锁定机制,在固定预算下提升多数票准确性并节省调用成本。
AI 中文摘要
对多个LLM响应进行投票是测试时扩展和集成推理中的常见原语。收集更多响应可以扩大候选池,并增加发现正确答案的机会。在固定的调用预算下,已发现的答案仍需在剩余调用中积累足够的支持,以成为最终的多数票获胜者,从而产生发现到决策的差距。在这项工作中,我们通过实现的投票状态和剩余调用预算来刻画这一差距。我们推导出一个尖锐的可恢复性阈值,并表明随着采样的进行,观察到的候选集只能扩大,而可达的端点获胜者集只能收缩,从而产生候选级别的转换窗口。在指定的独立同分布响应定律下,相同的状态产生精确的有限时域端点概率。我们进一步表明,合并错误答案的身份保留了单次调用的正确性,并且不能提高多数票的准确性,而重新分配错误答案概率的效果取决于实现的投票状态。单例可达性产生一个无金标准的精确锁定证书。对于已知的答案宇宙,其首次触发是最早的前缀,在该前缀处所有允许的延续都产生相同的固定预算输出。经验上,大多数被发现但未被选中的正确答案仅在发现后失去可达性。在受控的Word16研究中,输入排列将原始多数票准确性提高了21.1个百分点,而单次调用的正确性基本不变。精确锁定在16次调用的预算下节省了28-30%的调用,同时保留了每个固定预算的输出。
英文摘要
Voting over multiple LLM responses is a common primitive in test-time scaling and ensemble inference. Collecting more responses can expand the candidate pool and increase the chance that a correct answer is discovered. Under a fixed call budget, a discovered answer still needs to accumulate enough support within the remaining calls to become the final plurality winner, creating a discovery-to-decision gap. In this work, we characterize this gap through the realized vote state and remaining call budget. We derive a sharp recoverability threshold and show that, as sampling proceeds, the observed candidate set can only expand while the set of reachable endpoint winners can only contract, inducing a candidate-level conversion window. Under a specified iid response law, the same state yields exact finite-horizon endpoint probabilities. We further show that merging wrong-answer identities preserves single-call correctness and cannot improve plurality accuracy, and that the effect of redistributing wrong-answer probability depends on the realized vote state. Singleton reachability yields a gold-free exact locking certificate. For a known answer universe, its first trigger is the earliest prefix at which all admissible continuations yield the same fixed-budget output. Empirically, most discovered-but-unselected correct answers lose reachability only after discovery. In a controlled Word16 study, input permutation improves raw-plurality accuracy by 21.1 points with essentially unchanged single-call correctness. Exact locking saves 28-30% of calls at a 16-call budget while preserving every fixed-budget output.