发表机构
University of Science and Technology of China; Nanjing University; Hefei University of Technology; Tsinghua University(中国科学技术大学; 南京大学; 合肥工业大学; 清华大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对流式化身生成中自强制蒸馏的动态崩溃问题,提出DynaForcing框架,通过三种策略及优化技术恢复动态、提升视觉质量,解决了质量与动态的权衡问题。
AI 中文摘要
音频驱动的化身生成需要逼真的唇同步、富有表现力的动作以及实时流式传输。近期研究通过结合分布匹配蒸馏(DMD)的自强制技术实现了实时流式传输,但该范式存在一个尚未被系统表征的关键缺陷:动态崩溃,即学生模型收敛至接近静态的最优解,虽感知质量较高但时间动态被严重抑制。我们将此问题归因于两个原因:DMD中的反向KL目标偏向低动态模式,以及无锚定的自条件化形成放大崩溃的反馈循环。这对化身生成危害极大,因为即使微小的动作损失也会破坏唇同步与表情。为解决该问题,我们提出DynaForcing,一种在不同层面应用三种互补策略的训练框架:数据层面的混合强制(Hybrid Forcing)将滚动输出锚定到真实动态以打破反馈循环;损失层面的动态感知奖励正则化(Dynamics-Aware Reward Regularization)通过对DMD的强化学习解释引入显式动作奖励,抵消反向KL偏差;条件层面的参考扰动(Reference Perturbation)扰动参考图像以解耦身份与静态细节,迫使模型依赖音频生成动作。我们还引入计算图剪枝与梯度重放,将自强制的GPU内存占用降低一个数量级以上。实验表明,DynaForcing可将动态恢复至与教师模型相当的水平(动态度Dyn-Deg:0.31→0.73,同步度Sync-C:7.03→7.68),同时提升视觉质量,在整个训练过程中解决了质量与动态之间的权衡问题,无需提前停止训练。
英文摘要
Audio-driven avatar generation requires realistic lip-sync, expressive motion, and real-time streaming. Recent work achieves the latter via self-forcing with Distribution Matching Distillation (DMD), but this paradigm suffers from a critical failure that has not been systematically characterized: dynamic collapse, where the student model converges to a near-static optimum with high perceptual quality but severely suppressed temporal dynamics. We trace this to two causes: the reverse KL objective in DMD, which biases toward low-motion modes, and unanchored self-conditioning, which creates a feedback loop that amplifies collapse. This is especially harmful for avatars, where even subtle motion loss breaks lip-sync and expression. To address this, we propose DynaForcing, a training framework with three complementary strategies applied at different levels. Specifically, Hybrid Forcing anchors rollouts to ground-truth dynamics at the data level to break the feedback loop. Dynamics-Aware Reward Regularization introduces explicit motion rewards via the RL interpretation of DMD to counteract the reverse KL bias at the loss level. Reference Perturbation perturbs reference images to decouple identity from static details, forcing the model to rely on audio for motion at the conditioning level. We further introduce computation graph pruning and gradient replay, reducing the GPU footprint of self-forcing by over an order of magnitude. Experiments show that DynaForcing recovers dynamics to teacher-comparable levels (Dyn-Deg: 0.31 -> 0.73, Sync-C: 7.03 -> 7.68) while improving visual quality, resolving the quality-dynamics trade-off throughout training without early stopping.
CommentsAccepted at ACM International Conference on Multimedia (MM '26)