arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

TACD:通过终端放大控制蒸馏高效文本到运动模型

TACD: Distilling Efficient Text-to-Motion Models via Terminal Amplification Control

Wei-Jin Huang, Yuan-Ming Li, Kun-Yu Lin, Wang Luo, Yinlin Zhu, Yue Yu, Shenghao Ye, Junbin Yuan, Fa-Ting Hong, Qing Zhang, Wei-Shi Zheng

arXiv 2610.02867首次发表:更新:

发表机构

Sun Yat-sen University; Nanyang Technological University; Wuhan University; University of Science and Technology of China; The Hong Kong University of Science and Technology(中山大学; 南洋理工大学; 武汉大学; 中国科学技术大学; 香港科技大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

TACD通过终端放大控制蒸馏,从文本和教师模型高效训练少步运动生成器,显著降低FID并实现数倍加速。

AI 中文摘要

近期的文本到运动模型提高了运动质量和指令遵循能力,但多步去噪和大型模型组件使得部署缓慢且内存密集。我们提出了终端放大控制蒸馏(TACD),一种从文本提示和预训练教师模型训练高效运动生成器的在线策略方法,无需真实运动训练数据。基于分段在线流蒸馏,我们监督学生生成轨迹上的干净运动预测。我们识别出一种失败模式,即在固定监督网格上的速度匹配反复过度加权去噪端点附近的误差,降低了少步生成的质量。TACD将最新的教师查询与学生步长绑定,限制了干净运动空间中有效损失权重的范围,而不改变推理过程。在HumanML3D和KIT-ML上的实验表明,少步生成得到改进,包括相对于无此约束的蒸馏,八步HY-Motion学生FID降低了58%。对于扩散教师,TACD的端点匹配形式在HumanML3D上产生了四步学生,其FID更低,文本-运动检索与50步教师相当或更好。在HY-Motion和Kimodo上,具有紧凑组件的八步学生实现了7.7-11.9倍的端到端加速,并将峰值GPU内存相对于教师减少了3.8-6.7倍。项目页面:此https URL

英文摘要

Recent text-to-motion models have improved motion quality and instruction following, yet many-step denoising and large model components make deployment slow and memory-intensive. We present Terminal-Amplification-Controlled Distillation (TACD), an on-policy approach for training efficient motion generators from text prompts and pretrained teachers, without real-motion training data. Building on segmented on-policy flow distillation, we supervise clean-motion predictions along student-generated trajectories. We identify a failure mode in which velocity matching on a fixed supervision grid repeatedly overweights errors near the denoising endpoint, degrading few-step generation. TACD ties the latest teacher query to the student's step size, bounding the effective loss weights in clean-motion space without changing inference. Experiments on HumanML3D and KIT-ML demonstrate improved few-step generation, including a 58% reduction in eight-step HY-Motion student FID relative to distillation without this bound. For diffusion teachers, the endpoint-matching form of TACD yields four-step students with lower FID and matched or improved text-motion retrieval relative to their 50-step teachers on HumanML3D. On HY-Motion and Kimodo, eight-step students with compact components achieve 7.7-11.9x end-to-end speedups and reduce peak GPU memory by 3.8-6.7x relative to their teachers. Project page: https://vkgo.github.io/TACD/

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑