扩展视频生成用于推理:代价是什么?
Scaling Video Generation for Reasoning: At What Cost?
浏览论文内容
中文总结 AI 辅助
本研究通过魔方预测基准,探究视频生成扩展对推理能力的影响及计算代价,发现小模型在有限计算下更优,且符号监督可提升状态准确率。
中文摘要 AI 辅助
我们研究扩展视频生成是否能使模型从过去的帧中推理出隐藏信息,以及其计算代价是多少。我们的受控基准要求从三个面的固定视角预测一个初始已解开的2x2x2魔方的九个规定动作。正确的预测需要推断动作如何改变隐藏状态,模拟器为评估提供了精确的真实值。模型早期学习到合理的魔方几何形状,而正确的贴纸配置则需要更多的训练。尽管验证均方误差遵循近似幂律缩放,但较低的均方误差损失并不能可靠地指示下游推理能力。较小的自回归模型在有限计算下达到更高的状态准确率,而较大的模型在更多训练后达到更高的准确率。在大约0.1 PF-days时,70M参数模型在44.6%的动作后帧中正确预测可见贴纸配置,而1B模型仅为0.3%,后者在3.14 PF-days时达到83.7%。在相同训练数据预算下,符号状态监督将20M模型的帧准确率从31.1%提升至67.3%,这表明学习状态变化的表示可以补充扩展。
英文摘要
We study whether scaling video generation enables models to reason about hidden information from the past frames, and at what computational cost. Our controlled benchmark requires predicting nine prescribed moves of an initially solved 2x2x2 Rubik's Cube from a fixed view of three faces. Correct predictions require inferring how actions change hidden states, and the simulator provides exact ground truth for evaluation. Models learn plausible cube geometry early, while correct sticker configurations require substantially more training. Although validation MSE follows approximate power-law scaling, lower MSE loss does not reliably indicate downstream reasoning capabilities. Smaller autoregressive models achieve higher state accuracy with limited compute, while larger models reach higher accuracy after more training. At roughly 0.1 PF-days, the 70M-parameter model correctly predicts the visible sticker configuration in 44.6% of post-action frames, compared with 0.3% for the 1B model, which reaches 83.7% at 3.14 PF-days. Symbolic state supervision raises the 20M model's frame accuracy from 31.1% to 67.3% at the same training-data budget, suggesting that learning representations of state changes can complement scaling.
发表机构
- Rice University(莱斯大学)
- Carnegie Mellon University(卡内基梅隆大学)
机构由 AI 辅助整理,请以论文原文为准。