发表机构
Hainan University; Xiamen University; Zhejiang Normal University(海南大学; 厦门大学; 浙江师范大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出无训练框架TRAC,通过稳健累积调度、轨迹感知引导调度和谱结构校正,解决自回归视频生成中误差累积问题,在SkyReels-V2和FramePack-F1上实现最高推理效率与最佳生成质量。
AI 中文摘要
在本文中,我们提出了轨迹感知复用与自适应校正(TRAC),这是一种用于高效自回归(AR)视频生成的无训练框架。现有的加速方法主要针对具有双向注意力的单轨迹生成。相比之下,AR视频生成按顺序耦合块级去噪轨迹。因此,近似误差会在生成过程中累积并传播。TRAC通过三个组件解决这一挑战,包括稳健累积调度(RCS)、自回归轨迹感知引导调度(ATGS)和谱结构校正(SSC)。RCS通过累积展开误差和跨块/提示变化来选择缓存复用调度。ATGS沿全局AR轨迹协调CFG刷新。SSC恢复第一个块的低频结构,以纠正长期结构损失。在SkyReels-V2和FramePack-F1上的实验表明,与现有方法相比,TRAC在AR视频生成中实现了最高的推理效率和最佳的生成质量。
英文摘要
In this paper, we present trajectory-aware reuse and adaptive correction (TRAC), a training-free framework for efficient autoregressive (AR) video generation. Existing acceleration methods mainly target single-trajectory generation with bidirectional attention. AR video generation, by contrast, sequentially couples chunk-level denoising trajectories. Consequently, approximation errors accumulate and propagate through the generation process. TRAC addresses this challenge with three components, including robust cumulative scheduling (RCS), autoregressive trajectory-aware guidance scheduling (ATGS), and spectral structure correction (SSC). RCS selects cache reuse schedules by cumulative rollout error and cross-chunk/prompt variation. ATGS coordinates CFG refreshes along the global AR trajectory. SSC restores low-frequency structure of the first chunk to correct long-term structural loss. Experiments on SkyReels-V2 and FramePack-F1 show that, compared with existing methods, TRAC achieves both the highest inference efficiency and the best generation quality for AR video generation.
CommentsPreprint under review