发表机构
Tsinghua University; Joy Future Academy, JD; Southeast University(清华大学; 京东探索研究院; 东南大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对流式虚拟形象生成中DMD蒸馏导致动态和多样性坍缩的问题,提出路由强制方法,按语义区域和噪声阶段路由蒸馏目标,动态性提升45%,多样性提升7-25%。
AI 中文摘要
音频驱动的流式虚拟形象生成需要实时合成与语音同步且具有动态和多样运动的视频。自强制(Self Forcing)利用分布匹配蒸馏(DMD)将双向视频扩散模型蒸馏为因果的、少步生成器,以实现实时流式生成。然而,DMD最小化反向KL散度,这本质上是模式寻找的:它导致学生模型丢弃高动态模式并坍缩到静态输出,压缩了生成视频的动态性和多样性。我们发现这种坍缩是区域异质性的:涉及姿态和手势变化的身体区域遭受最大的多样性损失,音频驱动的嘴部区域损失较小,而背景几乎保持稳定。基于这一观察,我们提出路由强制(Routed Forcing),它通过语义区域和噪声阶段路由蒸馏目标,以改善动态性和多样性,同时保持视觉质量。具体而言,(1)何处强制:语义区域路由将数据强制蒸馏(DFD)应用于多样性坍缩最严重的身体区域,DFD使用真实视频监督学生模型,而对嘴部和背景保留DMD以保持唇部同步和场景稳定性。(2)何时强制:噪声阶段路由在高噪声阶段激活DFD,此时真实视频作为有效监督以注入多样和动态的运动模式。在低噪声阶段,使用DMD细化细节,避免真实视频与学生生成视频之间的空间差异导致的模糊和伪影。实验表明,与自强制相比,路由强制将动态性提升高达45%,多样性提升7-25%,同时保持视频质量和唇部同步。
英文摘要
Audio-driven streaming avatar generation requires real-time synthesis of speech-synchronized videos with dynamic and diverse motion. Self Forcing uses Distribution Matching Distillation (DMD) to distill bidirectional video diffusion models into causal, few-step generators for real-time streaming. However, DMD minimizes a reverse KL divergence, which is inherently mode-seeking: it causes the student to discard high-dynamic modes and collapse onto static outputs, compressing both dynamics and diversity of generated videos. We find that this collapse is region-heterogeneous: person regions involving pose and gesture variations suffer the largest diversity loss, the audio-driven mouth region shows a small loss, and the background remains nearly stable. Based on this observation, we propose Routed Forcing, which routes the distillation objective by semantic region and noise stage to improve dynamics and diversity while preserving visual quality. Specifically, (1) Where to Force: Semantic-Region Routing applies Data-Forcing Distillation (DFD), which supervises the student with real videos, to the person region where diversity collapse is most severe, while retaining DMD for the mouth and background to preserve lip synchronization and scene stability. (2) When to Force: Noise-Stage Routing activates DFD at high noise stages, where real video serves as effective supervision to inject diverse and dynamic motion patterns. At low noise stages, DMD is used to refine details, avoiding blur and artifacts from spatial differences between real video and student-generated video. Experiments show that Routed Forcing improves dynamics by up to 45% and diversity by 7-25% over Self Forcing, while preserving video quality and lip synchronization.