解耦自强制蒸馏用于流式说话头生成
Decoupled Self-Forcing Distillation for Streaming Talking Head Generation
浏览论文内容
中文总结 AI 辅助
提出解耦自强制蒸馏,在身份解耦运动空间融合条件,用小型因果自回归变换器生成运动并由扩散渲染器转视频,实现流式说话头生成,达15.4 FPS、1.3秒延迟且保真度不降。
中文摘要 AI 辅助
流式说话头生成在驱动音频到达时逐帧生成画面,然而保真度和效率迄今相互制约:端到端方法将视频扩散模型直接以音频为条件,实现高质量但仅在大规模下可行,而较廉价的两阶段方法生成中间运动表示,在保真度上落后。我们认为前者的成本在于融合目标:视频潜变量由身份、外观和背景主导,这些均与音频无关,因此将音频耦合到每个像素会模糊细节并浪费容量。我们转而将条件融合在低维的身份解耦运动空间中,按时间粒度路由音频和运动描述,并使用一个小型因果自回归变换器生成运动潜变量,由预训练的扩散渲染器将其转为视频。条件因此间接控制视频,高保真不再需要大型骨干网络。对此分解进行流式化需要两个模型均为因果的,而暴露偏差问题在给定双向教师的情况下可通过自强制解决,但运动空间中不存在这样的教师。我们的解耦自强制蒸馏在单个冻结教师下解决两个模型:以运动为条件时,将渲染器蒸馏为块因果学生;无条件时,对渲染出的展开序列与真实视频评分,通过其生成的视频监督运动。这将保真度上限从运动生成器提升至更强的渲染器。两个模型作为并行因果流运行,在1.3秒延迟下达到15.4 FPS,且无质量下降。
英文摘要
Streaming talking-head generation produces each frame as its driving audio arrives, yet fidelity and efficiency have so far pulled in opposite directions: end-to-end methods condition a video diffusion model on audio directly and achieve high quality but only at large scale, while cheaper two-stage methods generate an intermediate motion representation and trail in fidelity. We argue the cost of the former lies in the target of fusion: the video latent is dominated by identity, appearance and background, none of which audio bears on, so coupling audio to every pixel blurs detail and wastes capacity. We instead fuse conditions in a low-dimensional identity-disentangled motion space, routing audio and motion captions by their temporal granularity, and generate motion latents with a small causal autoregressive transformer that a pretrained diffusion renderer turns into video. Conditions thus control video transitively, and high fidelity no longer requires a large backbone. Streaming this decomposition needs both models to be causal, and the exposure-bias problem could be solved by self-forcing given a bidirectional teacher. But there is no such teacher in motion space. Our decoupled self-forcing distillation resolves both models under one frozen teacher: conditioned on motion, it distills the renderer into a block-causal student; unconditionally, it scores rendered rollouts against real videos, supervising motion by the video it produces. This lifts the fidelity ceiling from the motion generator onto the stronger renderer. The two models run as parallel causal streams, reaching 15.4 FPS at 1.3 s latency with no quality degradation.
发表机构
- Shanghai Jiao Tong University(上海交通大学)
- Ant Group(蚂蚁集团)
机构由 AI 辅助整理,请以论文原文为准。