发表机构
TaoLive AIGC, Taobao & Tmall Group of Alibaba(阿里巴巴淘宝天猫集团淘Live AIGC)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
TBDub通过任务自适应后训练和少步蒸馏改进X-Dub,提升生产环境下的视觉配音鲁棒性、稳定性和效率,教师模型在多项指标上优于X-Dub,学生模型大幅加速且保持质量。
AI 中文摘要
视觉配音必须在替换语音的同时同步嘴部运动,并保持身份、外观和时间一致性。尽管X-Dub提供了一个强大的无掩码视频编辑基线,但其在直播和生成视频内容上的应用暴露了在生产领域鲁棒性、时间和运动稳定性、身份和口腔细节保留以及推理效率方面的局限性。我们提出了TBDub,一个面向生产的X-Dub扩展,结合了任务自适应后训练和任务感知的少步蒸馏。后训练使用生产领域数据、生产特定条件和过滤以及增强的音频特征来适配视频DiT,以获得30步的教师模型。蒸馏将DMD/DMD2适配到条件视频编辑,并将教师模型压缩为两步的学生模型。在38个TalkVid片段上,教师模型在所有八项报告的重建、感知、身份和同步指标上均优于X-Dub。在MOS评估中,教师模型在唇形同步一致性、身份一致性和视觉质量上分别比X-Dub提高了0.14、0.95和0.90分,而学生模型在唇形同步和视觉质量上得分最高,并在身份一致性上接近教师模型。在单个NVIDIA H20 GPU上,从第一次VAE编码到最终VAE解码的成对端到端生成计时中,在512×512分辨率下,学生模型达到7.13有效FPS,并将总延迟降低了13.93倍;仅DiT阶段就加速了42.49倍。学生模型在很大程度上保留了教师模型的生成质量和视听同步。代码可在GitHub上获取,30步教师模型和两步学生模型的权重可在Hugging Face上获取。
英文摘要
Visual dubbing must synchronize mouth motion with replacement speech while preserving identity, appearance, and temporal consistency. Although X-Dub provides a strong mask-free video-editing baseline, its application to livestream and generated-video content reveals limitations in production-domain robustness, temporal and motion stability, identity and oral-detail preservation, and inference efficiency. We present \textbf{TBDub}, a production-oriented extension of X-Dub that combines task-adaptive post-training with task-aware few-step distillation. Post-training adapts the video DiT using production-domain data, production-specific conditioning and filtering, and enhanced audio features to obtain a 30-step Teacher. Distillation adapts DMD/DMD2 to conditional video editing and compresses the Teacher into a two-step Student. On 38 TalkVid clips, the Teacher improves all eight reported reconstruction, perceptual, identity, and synchronization metrics over X-Dub. In the MOS evaluation, it improves lip-sync consistency, identity consistency, and visual quality over X-Dub by 0.14, 0.95, and 0.90 points, while the Student achieves the highest lip-sync and visual-quality scores and remains close to the Teacher in identity consistency. In paired end-to-end generation timing from the first VAE encode through the final VAE decode on a single NVIDIA H20 GPU at $512\times512$, the Student reaches 7.13 effective FPS and reduces total latency by $13.93\times$; the DiT stage alone is accelerated by $42.49\times$. The Student largely retains the Teacher's generation quality and audiovisual synchronization. The code is available on GitHub at \https://github.com/TaoLiveAIGC/TBDub, and the 30-step Teacher and two-step Student weights are available on Hugging Face at https://huggingface.co/TaoLiveAIGC/TBDub.