发表机构
Apple(苹果公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出后训练框架IVT,在训练时内化视觉预测能力,推理时直接生成答案,相较Visual CoT性能相当或更优且延迟降低5倍以上,可实现更高效准确的主动视频推理。
AI 中文摘要
多模态大语言模型越来越多地使用视觉思维链(Visual CoT)对空间、时间和具身环境进行推理。通过生成中间推理图像,Visual CoT提供了一种直观的视觉预见机制,但会引入大量推理开销,这对于主动视频推理而言尤其成问题。我们探究模型是否能在训练时学习进行视觉思考,同时在推理时直接推理。我们提出了内化视觉思维(Internalized Visual Thinking, IVT),这是一种后训练框架,它在未标记视频上联合优化文本预测和下一嵌入预测。给定部分观测的视频,IVT预测未来帧的潜在表示以及目标文本答案,促使模型捕捉运动、物体转换、交互和潜在意图。在推理时,IVT直接生成答案,无需合成或重新编码未来帧。我们对目标表示、解码器设计、预测范围、数据混合、训练课程和预测目标进行了对照研究。IVT在全部六个评估设置中均优于直接答案微调,同时保留相同的推理路径。与显式Visual CoT相比,IVT达到了相当或更好的性能,并将平均端到端延迟降低了5倍以上。我们的研究结果表明,在推理时进行显式像素空间生成(如视觉思维链所用)对于有效的主动视频推理而言可能并非必要。预测世界建模可在训练时内化,以生成更准确且效率显著更高的多模态推理器。
英文摘要
Multimodal large language models increasingly use visual chain-of-thought (Visual CoT) to reason about spatial, temporal, and embodied environments. By generating intermediate reasoning images, Visual CoT provides an intuitive mechanism for visual foresight but introduces substantial inference overhead, which is particularly problematic for proactive video reasoning. We ask whether models can learn to think visually during training while reasoning directly at inference. We introduce Internalized Visual Thinking (IVT), a post-training framework that jointly optimizes textual prediction and next-embedding prediction over unlabeled videos. Given a partially observed video, IVT predicts latent representations of future frames together with the target textual answer, encouraging the model to capture motion, object transitions, interactions, and latent intent. At inference, IVT generates the answer directly without synthesizing or re-encoding future frames. We conduct controlled studies across target representations, decoder designs, prediction horizons, data mixtures, training curricula, and predictive objectives. IVT improves over direct-answer fine-tuning on all six evaluation settings while retaining the same inference pathway. Compared with explicit Visual CoT, IVT achieves comparable or better performance and reduces average end-to-end latency by more than 5x. Together, our findings suggest that explicit pixel-space generation at inference time, as used in visual chain-of-thought, may not be necessary for effective proactive video reasoning. Predictive world modeling can be internalized during training to produce multimodal reasoners that are both more accurate and substantially more efficient.