arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.13257cs.CVcs.AI

采样余量并非选择增益:视频世界模型测试时扩展的计算价值审计

Sampling headroom is not selection gain: a compute-value audit of test-time scaling for video world models

  • Tsinghua University(清华大学)
  • University of Technology Sydney(悉尼科技大学)

机构由 AI 辅助整理,请以论文原文为准。

Yuhua Jiang, Junjie Lu, Feifei Gao

AI总结:

本研究通过计算价值审计框架,证明视频世界模型中采样余量不等于选择增益,仅当额外采样能转化为可靠决策且收益超过计算成本时才有价值。

AI中文摘要:

测试时扩展(TTS)只有在额外计算产生更好的候选且系统能够可靠地识别它们时,才能改进生成。这一区分对于视频世界模型尤为重要,因为更宽的样本池可能包含更强的推演,却不会改进最终被选择的输出。我们引入了计算价值审计(CVA),这是一个顺序框架,询问额外采样是否创造机会、可观察信号是否提供可靠状态、该状态是否支持有益行动,以及由此产生的增益是否超过生成和验证的全部入场费用。在192个Physics-IQ场景中,将样本池从4个扩展到16个,使oracle质量提高了+9.23 IQ(95%置信区间[+7.44, +11.14]),但Flow、Cycle和VideoReward未能可靠地恢复这一余量。在三个生成器中,十二种自适应深度策略均未优于均匀计算;它们仅恢复了所测入场费用的42-69%。在VideoPhy2上进行的匹配60-NFE的预测-扰动干预,在三个新种子重复实验中同样为负。这些负面结果并非普遍:anchor-explorer在稀疏PRM800K设置中通过了所有四个阶段,MMLU-Pro暴露了预测状态与有用行动之间的差距,而特权配对未来建立了正面的视频上限。总之,这些结果表明,采样余量仅当能被转化为可靠决策且其收益能承受完整计算费用时,才具有部署价值。代码可在该https URL获取。

英文摘要:

Test-time scaling (TTS) can improve generation only when additional compute produces better candidates and the system can reliably identify them. This distinction is especially important for video world models, where a wider sample pool may contain stronger rollouts without improving the output that is ultimately selected. We introduce the Compute-Value Audit (CVA), a sequential framework that asks whether extra sampling creates opportunity, observable signals provide a reliable state, that state supports a beneficial action, and the resulting gain exceeds the full entry fee of generation and verification. On 192 Physics-IQ scenes, expanding the pool from 4 to 16 candidates increases oracle quality by +9.23 IQ (95% CI [+7.44, +11.14]), but Flow, Cycle, and VideoReward fail to recover this headroom reliably. Across three generators, none of twelve adaptive-depth policies outperforms uniform compute; they recover only 42-69% of the measured entry fee. A matched-60-NFE Predict-and-Perturb intervention on VideoPhy2 is likewise negative across three fresh-seed replicas. These negative results are not universal: anchor-explorer passes all four stages in a sparse PRM800K setting, MMLU-Pro exposes the gap between predictive state and useful action, and a privileged paired future establishes a positive video upper bound. Together, these results show that sampling headroom has deployment value only when it can be converted into a reliable decision whose benefit survives the complete compute charge. Code is available at https://github.com/YuhuaJiang2002/sampling-headroom-is-not-selection-gain.

补充信息

↑