arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.02914cs.CV

自定义强制:自回归视频生成的无训练主体定制

Custom Forcing: Training-Free Subject Customization for Autoregressive Video Generation

Yunseung Ok, Hyunsoo Kim, Minseo Kim, Suhyun Kim

首次发表
浏览论文内容

中文总结 AI 辅助

针对自回归视频生成中主体定制问题,提出无训练的自定义强制方法,通过持久KV缓存锚定帧并采用漂移自适应值放大和锚定对比引导,在保持身份和速度上优于现有方法。

中文摘要 AI 辅助

自回归视频模型能够实时生成分钟级视频,但它们生成的是来自文本的通用主体,而非用户提供图像中的特定主体。现有的定制方法要么需要昂贵的逐主体优化,要么使用预训练的条件网络,该网络以双向注意力联合处理所有视频帧。这两种方法都不是为因果流式处理设计的。我们提出了自定义强制(Custom Forcing),一种无训练方法,将基于参考的锚定帧存储在冻结的自回归视频模型的持久KV缓存中。然而,固定锚定面临两个限制:简单的条件化会导致身份漂移,且文本提示继续偏向通用主体。为解决这些问题,漂移自适应值放大(DVA)根据身份漂移程度缩放参考影响,而锚定对比引导(ACG)则引导生成远离通用类别先验。在两分钟以上的生成过程中,固定锚定的DINO-I从0.58降至0.42,而自定义强制将其保持在0.58至0.62之间且不降低运动质量。自定义强制还实现了比双向定制方法更高的主体相似度,并在30秒内比因果图像到视频和参考到视频模型更好地保持身份,同时每帧生成速度比这些长视频基线快9.5至28.5倍。

英文摘要

Autoregressive video models can generate minute-long videos in real time, but they produce generic subjects from text rather than specific subjects from user-provided images. Existing customization methods either require costly per-subject optimization or use pretrained conditioning networks that jointly process all video frames with bidirectional attention. Neither approach is designed for causal streaming. We present Custom Forcing, a training-free method that stores reference-based anchor frames in the persistent KV cache of a frozen autoregressive video model. However, fixed anchors face two limitations: simple conditioning allows identity to drift, and the text prompt continues to favor a generic subject. To address these problems, drift-adaptive value amplification (DVA) scales reference influence with the degree of identity drift, while anchor contrast guidance (ACG) steers generation away from the generic class prior. Over two-minute rollouts, fixed anchors fall from 0.58 to 0.42 in DINO-I, while Custom Forcing keeps it between 0.58 and 0.62 without reducing motion. Custom Forcing also achieves higher subject similarity than bidirectional customization methods and better preserves identity over 30s than causal image-to-video and reference-to-video models, while generating each frame 9.5-28.5 times faster than these long-video baselines.

发表机构

  • Kyung Hee University(庆熙大学)
  • The University of Texas at Austin(德克萨斯大学奥斯汀分校)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

相关深度报道

↑