arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

CtrlCache:利用控制感知缓存加速交互式视频世界模型

CtrlCache: Accelerating Interactive Video World Models with Control-Aware Caching

Shangye Song, Dong Gong, Hong Jia, Yun Sing Koh, Xinyu Zhang

arXiv 2610.08777首次发表:更新:

发表机构

University of Auckland; UNSW Sydney(奥克兰大学; 新南威尔士大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出CtrlCache,一种免训练的控制感知缓存框架,利用控制序列调度计算,在交互式视频世界模型中实现1.21-1.41倍加速并提升生成质量。

AI 中文摘要

交互式视频世界模型需要高效地生成每个视频块,同时忠实地响应用户控制。许多系统采用分块自回归生成和少步去噪,但每个块仍需要数次昂贵的去噪迭代。免训练缓存可以减少这种成本,但现有策略主要基于模型内部的去噪动态做出重用决策,并未明确考虑控制转换。实际上,交互式生成明确暴露了一个他们未使用的信号:一个块的控制在其去噪之前到达,因此由它们派生的调度不会产生额外的前向传播成本。为此,我们分析了不同控制机制下的相邻块,发现结构相似性在动作变化时下降,而低频结构比高频细节更持久。基于这些观察,我们提出了CtrlCache,一个免训练的控制感知缓存框架,使计算适应当前控制序列。具体而言,动作感知调度和刷新策略检测块内和跨块的动作变化,并将每个块标记为初始、转换、转向或稳定状态。在一个选定的内部去噪步骤中,初始和转换块保留完整计算,而转向和稳定块重用同一块中最近完全计算步骤的变换器残差。为了利用稳定交互期间低频结构的持久性,我们进一步引入了频率混合历史先验引导,该引导在不额外进行DiT前向传播的情况下,整合了来自先前干净潜变量的补充信息。在Matrix-Game 2.0和LingBot-World v1/v2上评估,CtrlCache在不重新训练模型的情况下实现了1.21倍至1.41倍的DiT骨干加速,同时提高了所有三个模型在WBench Overall分数上的原始推理性能。

英文摘要

Interactive video world models need to generate each video chunk efficiently while responding faithfully to user controls. Many systems use chunk-wise autoregressive generation with few-step denoising, but each chunk still requires several costly denoising iterations. Training-free caching can reduce this cost, yet existing policies make reuse decisions primarily from model-internal denoising dynamics and do not explicitly account for control transitions. Actually, interactive generation explicitly exposes a signal they do not use: the controls for a chunk arrive before it is denoised, so a schedule derived from them costs no forward pass. To this end, we analyze adjacent chunks under different control regimes and find that structural similarity drops around action changes, while low-frequency structure remains more persistent than high-frequency detail. Motivated by these observations, we propose CtrlCache, a training-free control-aware caching framework that adapts computation to the current control sequence. Specifically, the action-aware scheduling and refresh policy detects action changes across and within chunks, and labels each chunk as initial, transition, turning, or steady state. At one selected interior denoising step, initial and transition chunks retain full computation, while turning and steady chunks reuse the transformer residual from the most recent fully computed step in the same chunk. To exploit the persistence of low-frequency structure during steady interaction, we further introduce a frequency-mixed history prior guidance that incorporates complementary information from the preceding clean latent without an additional DiT forward pass. Evaluated on Matrix-Game 2.0 and LingBot-World v1/v2, CtrlCache achieves 1.21x to 1.41x DiT-backbone speedups without model retraining while improving WBench Overall scores over original inference across all three models.

Comments18 pages. Project page: https://wrecklong.github.io/CtrlCache/

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑