arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于自一致性的扩散视频推理学习

Learning via Self-Consistency for Diffusion-based Video Reasoning

Zhenghao Ni, Weimin Qiu, Meng Tang

arXiv 2609.36826首次发表:更新:

发表机构

University of Toronto; University of California, Merced(多伦多大学; 加州大学默塞德分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出利用自一致性改进扩散视频推理,通过多轨迹聚合共识预测并蒸馏至模型,显著提升视觉搜索与迷宫求解等任务准确率。

AI 中文摘要

视频生成模型已展现出在视觉推理、感知及其他视觉任务中涌现的零样本能力。然而,基于扩散的视频生成本质上是随机的,而许多下游视觉任务是确定性的。受自一致性在大语言模型思维链推理中有效性的启发,我们研究了自一致性是否也能类似地改进基于扩散的视频推理。我们首先引入一种无需训练、测试时扩展的方法,该方法采样多个视频生成结果,并通过自一致性聚合其预测。具体而言,我们从多个生成轨迹中提取路径、位置或掩码,并将其聚合为共识预测。为了减少多次生成推理的开销,我们在去噪轨迹早期读取预测,这既保持了共识质量,又将去噪步数减少了一半以上。我们进一步提出拒绝微调(RFT),将共识预测蒸馏到视频生成模型中。所得模型内化了多样本共识的优势,推理时仅需单次生成,同时显著优于原始模型。在迷宫求解、视觉搜索和指代分割三项任务上的实验表明,我们的自一致性推理和共识蒸馏均大幅提升了基于视频的感知与推理能力,且无需真实视频或任务特定验证。对于视觉搜索,自一致性将任务准确率从单次生成的48.4%提升至99.0%。蒸馏模型通过单次生成保留了大部分共识优势。对于4x4迷宫求解,基于共识的训练将单次生成的严格成功率从72.0%提升至84.0%,且推理延迟不变。

英文摘要

Video generation models have demonstrated emerging zero-shot capabilities for visual reasoning, perception, and other vision tasks. However, diffusion-based video generation is inherently stochastic, while many downstream vision tasks are deterministic. Motivated by the effectiveness of self-consistency in chain-of-thought reasoning for large language models, we investigate whether self-consistency can similarly improve diffusion-based video reasoning. We first introduce a training-free test-time scaling method that samples multiple video generations and aggregates their predictions through self-consistency. Specifically, we aggregate extracted paths, locations, or masks from multiple rollouts into a consensus prediction. To reduce the inference overhead of multi-rollout generation, we read out predictions early in the denoising trajectory, which preserves consensus quality while reducing denoising steps by more than half. We further propose Rejection Fine-Tuning (RFT) to distill consensus predictions into the video generation model. The resulting model internalizes the benefit of multi-sample consensus and requires only a single generation at inference time, while substantially outperforming the original model. Experiments on three tasks, including maze solving, visual search, and referring segmentation, show that both our self-consistency inference and consensus distillation dramatically improve video-based perception and reasoning, without requiring ground-truth videos or task-specific verification. For visual search, self-consistency raises task accuracy from 48.4% for a single generation to 99.0%. The distilled model retains much of the consensus benefit with a single rollout. For 4-by-4 maze solving, consensus-based training improves the single-generation strict success rate from 72.0% to 84.0% with the same inference latency.

Comments23 pages, including 10 pages for the main body

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑