arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.18931cs.CLcs.AI

野外测试时缩放:为何利用而非探索是瓶颈

Test-Time Scaling in the Wild: Why Exploitation, Not Exploration, Is the Bottleneck

Davide Romano, Kanak Raj, Jerrod Parker, Daniele Giofrè

首次发表
浏览论文内容

中文总结 AI 辅助

该研究通过首个计算归一化对比发现,测试时缩放(TTS)中从候选池选择的利用环节是瓶颈,跨候选合成(Fusion)仅恢复约40%可用质量,探索缩放有效但利用失效。

中文摘要 AI 辅助

测试时缩放(TTS)通过增加推理计算量来提升语言模型输出,具体方式包括生成多个候选、在部分序列上搜索或迭代优化草稿。这些技术在数学和代码任务上取得了显著收益,但几乎仅在验证直接明了的任务上开发和测试。我们基于统一框架,将各方法的token预算有效性分解为探索与利用,在涵盖医学、法律、金融、通用对话和创意写作的5个开放式生成基准上,对5类TTS方法进行了首个计算归一化的对比。结果取决于分解的哪一侧:探索缩放有效,所有场景下,候选池中的最优候选随计算量增加稳步提升;而失效的是利用环节,即从丰富候选池转换为最终输出的步骤。对于最先进的生成器,奖励模型与真实质量的相关性仅为ρᵥ≈0.12,无论预算如何,选择过程近乎随机;树搜索通过多样性崩溃放大了这一失效;优化仅在5个基准中的1个上有帮助,在其他基准上的表观收益是混淆因素导致的;仅跨候选合成(Fusion)始终优于单样本基线,但仍仅恢复约40%的可用质量。候选池并非瓶颈,从候选池中进行选择才是。

英文摘要

Test-time scaling (TTS) improves language model outputs by spending additional inference compute - generating multiple candidates, searching over partial sequences, or iteratively refining drafts. These techniques yield large gains on mathematics and code, but have been developed and stress-tested almost exclusively on tasks where verification is straightforward. We conduct the first compute-normalised comparison of five TTS families across five open-ended generation benchmarks spanning medicine, law, finance, general chat, and creative writing - grounded in a unified framework that decomposes the effectiveness of each method's token budget into exploration and exploitation. The answer depends on which side of that decomposition you examine. Scaling exploration works: the best candidate in the pool improves steadily with compute across all settings. What breaks is exploitation - the step that converts a rich candidate pool into a final output. With state-of-the-art generators, reward models correlate at only $ρ_v \approx 0.12$ with true quality, rendering selection near-random regardless of budget. Tree search amplifies this failure through diversity collapse. Refinement helps on one of five benchmarks; its apparent gains elsewhere are confounded. Only synthesis across candidates (Fusion) consistently improves over single-sample baselines, yet still recovers only ~40% of available quality. The candidate pool is not the bottleneck - choosing from it is.

发表机构

  • Thomson Reuters(汤姆森路透)

机构由 AI 辅助整理,请以论文原文为准。

↑