发表机构
XPeng; The Chinese University of Hong Kong; Tsinghua University(小鹏汽车; 香港中文大学; 清华大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出连接自强迫训练框架,通过捷径梯度重放重新连接自回归视频生成的跨块梯度路径,结合分布匹配蒸馏,在不改变推理流程的前提下提升了长时程视频的视觉质量与时间一致性。
AI 中文摘要
为了在流式传输长视频时保持视觉质量和时间一致性,自强迫(Self Forcing)通过在自生成历史上进行自展开训练并结合键值(KV)缓存来缓解暴露偏差。为了保持内存可管理性,它会分离历史缓存,保留块之间的前向依赖关系但切断后向梯度路径。我们提出连接自强迫(Connected Self Forcing),这是一种重新连接自回归块之间梯度路径的训练框架,允许来自后续预测的反馈指导早期上下文的生成。这些连接超越了历史KV写入:梯度通过生成的潜变量传递到产生它们的计算中,将早期上下文的生成与其在后续预测中的使用关联起来。为了使这种连接训练具有内存效率,我们开发了捷径梯度重放(shortcut gradient replay),无需保留完整的展开计算图即可恢复跨块梯度。结合分布匹配蒸馏,连接自强迫根据历史块的直接监督及其对后续生成的贡献来训练这些块。在自回归视频生成上的实验表明,该方法在不改变推理过程的情况下,提升了长时程视觉质量和时间一致性。
英文摘要
To stream long videos while maintaining visual quality and temporal consistency, Self Forcing mitigates exposure bias through self-rollout training on self-generated histories with key-value (KV) caching. To keep memory manageable, it detaches historical caches, preserving forward dependencies between chunks but severing the backward gradient paths. We introduce Connected Self Forcing, a training framework that reconnects gradient paths across autoregressive chunks, allowing feedback from later predictions to guide how earlier context is generated. These connections go beyond historical KV-writing: gradients pass through generated latents into the computations that produced them, linking the generation of earlier context to its use in later predictions. To make this connected training memory-efficient, we develop shortcut gradient replay, which recovers cross-chunk gradients without retaining the full rollout computation graph. Integrated with distribution matching distillation, Connected Self Forcing trains historical chunks according to both their direct supervision and their contribution to subsequent generation. Experiments on autoregressive video generation show improvements in long-horizon visual quality and temporal consistency, without changing the inference procedure.
CommentsProject Page: https://eastbeanzhang.github.io/CSF/