arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.30096cs.CVcs.AI

通过免训练轨迹路由加速视频扩散

Accelerating Video Diffusion via Training-Free Trajectory Routing

Mustafa Munir, Huy Vu, Shreyas Misra, Rohit Jena, Sajad Norouzi, Ali Taghibakhshi, Anis Ahmad, Anjul Patney, Pavlo Molchanov, Nima Tajbakhsh

中文总结 AI 辅助

TRACK通过校准分歧分数在去噪步骤间切换大小模型,免训练实现视频扩散加速,在多个模型上获得约2倍速度提升且保持质量。

中文摘要 AI 辅助

视频扩散在计算上代价高昂,因为它需要在许多去噪步骤中执行大型模型。即使采用步骤蒸馏,推理仍然昂贵,因为每个蒸馏步骤仍需要一次代价高昂的模型评估。我们提出TRACK:通过top-K选择进行轨迹感知容量路由,这是一种异构去噪策略,在选定的步骤中在兼容的大型和小型模型之间切换,从而降低每次去噪评估的平均成本。切换步骤通过校准过程确定。TRACK首先使用大型模型推出参考轨迹。然后在每一步,也收集小型模型的预测,并与大型模型的预测进行比较,以获得相对分歧分数。两个模型接收相同的潜变量、时间步、条件和引导输入。在校准集上聚合该信号,生成跨扩散步骤的分歧分数图,该图决定了高效推理过程的切换策略:对质量敏感的步骤继续使用大型模型,而分歧分数低的步骤则路由到小型模型。推理在每一步仅执行选定的模型,无需重新训练、架构或调度器更改,也无需在线双模型评估。在Wan 2.1、Cosmos 3、TurboDiffusion和FastVideo上,TRACK分别实现了1.95倍、2.04倍-2.73倍、2.69倍和2.17倍的加速,同时保持了可比的整体质量和较高的多样性保持。因此,TRACK将自动化的免训练模型切换确立为视频扩散的一种实用加速范式。

英文摘要

Video diffusion is computationally expensive, as it requires executing a large model across many denoising steps. Even with step-distillation, inference remains expensive because every distilled step still requires a costly model evaluation. We present TRACK: TRajectory-Aware Capacity routing via top-K selection, a heterogeneous denoising strategy that switches between compatible large and small models at selected steps, reducing the average cost per denoising evaluation. The switching steps are determined using a calibration process. TRACK first rolls out a reference trajectory with the large model. Then at each step, the small model's prediction is also collected and compared against the large model's prediction to obtain a relative disagreement score. Both models receive the same latent, timestep, conditioning, and guidance inputs. Aggregating this signal over a calibration set produces a disagreement score map across diffusion steps, which determines a switching policy for an efficient inference process: quality-sensitive steps keep using the large model, while steps with low disagreement scores are routed to the small model. Inference executes only the selected model at each step, requiring no retraining, architecture or scheduler changes, or online dual-model evaluation. Across Wan 2.1, Cosmos 3, TurboDiffusion, and FastVideo, TRACK yields $1.95\times$, $2.04\times$-$2.73\times$, $2.69\times$, and $2.17\times$ speedups, respectively, with comparable aggregate quality and high diversity retention. TRACK thereby establishes automated, training-free model switching as a practical acceleration paradigm for video diffusion.

↑