arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.18379cs.LGcs.AIcs.CL

选择、重组还是重新求解?一种无候选控制的单次测试时聚合方法

Selection, Recombination, or a Fresh Solve? A Candidate-Free Control for Single-Pass Test-Time Aggregation

  • ETH Zurich(苏黎世联邦理工学院)

机构由 AI 辅助整理,请以论文原文为准。

Guiv Farmanfarmaian

AI总结:

该研究针对测试时聚合提出无候选控制方法,在数学基准测试中探究候选条件作用对准确率的影响,发现多正确候选时提升、全错时降低,明确了其效果边界。

AI中文摘要:

当所有候选答案均错误时,正确候选选择不可用,但聚合调用仍可重新求解问题。因此,正确的聚合答案可能反映重组、重新求解或两者兼具。为实现高效的测试时推理,关键问题在于候选上下文是否能在额外生成轮次之外提供额外价值。我们在相同的最大输出 token 配额下引入了缺失的无候选控制方法,并按正确候选的数量进行分层。在使用 Qwen3-4B 的 AIME-2025 和 HMMT-2025 基准测试中,当存在多个正确候选时,候选条件作用可提升准确率(Δ_cand(c2+) = +0.290);当所有候选均错误时,候选条件作用会降低准确率(Δ_cand(c0) = -0.123);而在仅有一个正确候选的情况下,其效果仍不明确。c2+ 和 c0 的结论在对自适应双基准程序进行保守修正后依然成立。在该反事实设定下,此规模下全错候选池的恢复效果的解释发生反转:与重新求解相比,基于全错候选池的条件作用会降低准确率。原始格式匹配和安慰剂结果可对失败情况进行描述性表征,但未明确其机制。在另一项结构化干预中,显式答案字段可因果引导输出向其值靠近;掩码操作未带来可测量的准确率提升,且未建立其与原始格式的等价性。上述证据仅限于 Qwen3-4B 系列模型、两个数学基准、首答截断的候选片段以及单次提示聚合。

英文摘要:

When every candidate is wrong, correct-candidate selection is unavailable, yet the aggregation call can still solve the problem afresh. A correct aggregate answer may therefore reflect recombination, fresh solving, or both. For efficient test-time reasoning, the relevant question is whether candidate context adds value beyond the additional generation pass. We introduce the missing candidate-free control under the same maximum output-token allowance and stratify by the number of correct candidates. Across AIME-2025 and HMMT-2025 with Qwen3-4B, candidate conditioning improves accuracy when multiple candidates are correct ($Δ_{\mathrm{cand}}$(c2+) = +0.290), lowers accuracy when every candidate is wrong ($Δ_{\mathrm{cand}}$(c0) = -0.123), and remains unresolved in the one-correct regime. The c2+ and c0 conclusions survive a conservative correction for the adaptive two-benchmark procedure. Under this counterfactual, the interpretation of all-wrong recovery reverses at this scale: conditioning on an all-wrong candidate pool lowers accuracy relative to a fresh solve. Original-format matching and placebo results characterize the failures descriptively but leave their mechanism unresolved. Within a separate structured intervention, explicit answer fields causally steer outputs toward their values; masking yields no measurable accuracy improvement, and equivalence with the original format was not established. The evidence is limited to one Qwen3-4B family, two mathematics benchmarks, first-answer-truncated candidate fragments, and single-pass prompted aggregation.

补充信息

↑