发表机构
Qualcomm AI Research(高通人工智能研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对自回归视频生成长视界性能下降问题,提出近期性强制方法,通过时间响应偏置调整注意力,弥合训练-推理差距,实现无额外推理成本的最优长视界生成。
AI 中文摘要
自回归(AR)视频生成在长视界上性能下降,原因在于一个被忽视的训练-推理差异,我们称之为KV驱逐不匹配:模型在短片段上训练,其中所有上下文帧都驻留在KV缓存中,但在推理时,内存限制迫使远距离帧从KV缓存中被驱逐——移除了模型所依赖的上下文。我们不是通过上下文截断来模拟驱逐——这会丢弃模型仍然需要的时间信息并损害运动连贯性——而是保留上下文,同时逐步降低远距离帧的影响,使其最终驱逐变得可忽略。为引导此设计,我们引入位置响应$R( \Delta, \\, t_{\text{denoise}})$,一种基于扰动的敏感性度量,揭示上下文影响随时间距离急剧衰减,并随去噪步骤系统变化。受此分析启发,我们提出近期性强制,它应用非正、时间步相关的偏置,称为时间响应偏置(TRB),直接基于$R$对softmax前的注意力对数进行偏置,在不修改上下文长度或训练目标的情况下弥合训练-推理差距。我们进一步引入偏置注意力重参数化(BAR),一种精确重表述,将偏置移出softmax,使TRB成为零开销的标准FlashAttention调用。近期性强制在免训练模式和基于训练模式下均有效。在VBench和VBench-Long上的实验表明,在无额外推理成本下实现了最先进的长视界生成质量。
英文摘要
Autoregressive (AR) video generation degrades over long horizons due to an overlooked train-inference discrepancy we term KV eviction mismatch: models train on short clips where all context frames reside in the KV cache, but at inference, memory constraints force distant frames to be evicted from the KV cache - removing context the model was conditioned on. Rather than simulating eviction via context truncation - which discards temporal information the model still needs and degrades motion coherence - we keep the context but while progressively reducing the influence of distant frames, making their eventual eviction negligible. To guide this design, we introduce the positional response $R( Δ, \, t_{\text{denoise}})$, a perturbation-based sensitivity measure revealing that context influence decays steeply with temporal distance and varies systematically across denoising steps. Motivated by this analysis, we propose Recency Forcing, which applies a non-positive, timestep-dependent bias, termed Temporal Response Bias (TRB), on pre-softmax attention logits derived directly from $R$, closing the train-inference gap without modifying context length or training objectives. We further introduce Biased Attention Reparameterization (BAR), an exact reformulation that moves the bias outside the softmax, making TRB a standard FlashAttention call at zero overhead. Recency Forcing operates in both training-free mode and training-based mode. Experiments on VBench and VBench-Long demonstrate state-of-the-art long-horizon generation quality at no additional inference cost.