发表机构
Posts and Telecommunications Institute of Technology(邮电技术学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
MixiMotion通过离线集合蒸馏和非对称双向匹配,实现严格单步文本到动作生成,在保持高质量的同时大幅降低推理延迟,并取得了质量与效率的良好平衡。
AI 中文摘要
迭代式文本到动作生成能够产生高质量且语义对齐的动作,但需要多次网络评估,导致显著的推理延迟。我们提出了MixiMotion,一个基于离线集合蒸馏的严格单步文本到动作生成框架。不同于为每个文本提示蒸馏单个教师轨迹,MixiMotion构建了一个包含多个教师动作的离线库,并通过非对称双向匹配对齐教师和学生样本集。教师到学生的方向促进了对教师支持动作的覆盖,而学生到教师的方向则抑制了不支持的生成。我们进一步引入了可微的解码空间运动学监督,以在解码动作空间中用约束补充归一化表示匹配。在推理时,MixiMotion通过单次网络评估生成完整的动作序列,无需教师查询、迭代采样或候选排序。在ViMoGen上,MixiMotion达到了0.835的语义对齐分数,优于所评估的单步基线,并接近其50步HY-Motion-1.0-Lite教师的0.858分。在盲人评估中,MixiMotion获得了4.33的总体评分,而教师为4.50,同时优于所评估的单步/少步基线。同时,生成延迟从829.58毫秒降低到9.30毫秒,相当于89.2倍的加速。这些结果展示了严格单步文本到动作生成中有效的质量-效率权衡。
英文摘要
Iterative text-to-motion generation delivers high-quality and semantically aligned motions but requires multiple network evaluations, resulting in substantial inference latency. We present \textbf{MixiMotion}, a strict one-step text-to-motion generation framework based on offline set distillation. Instead of distilling a single teacher trajectory for each text prompt, MixiMotion constructs an offline bank of multiple teacher motions and aligns teacher and student sample sets through \textbf{asymmetric bidirectional matching}. The teacher-to-student direction promotes coverage of diverse teacher-supported motions, while the student-to-teacher direction suppresses unsupported generations. We further introduce differentiable decoded-space kinematic supervision to complement normalized representation matching with constraints in the decoded motion space. At inference, MixiMotion generates a complete motion sequence with a single network evaluation, without teacher queries, iterative sampling, or candidate ranking. On ViMoGen, MixiMotion achieves a semantic alignment score of $0.835$, outperforming the evaluated one-step baselines and approaching the $0.858$ score of its 50-step HY-Motion-1.0-Lite teacher. In blinded human evaluation, MixiMotion obtains an overall rating of $4.33$, compared with $4.50$ for the teacher, while outperforming the evaluated one-/few-step baselines. Meanwhile, generation latency is reduced from $829.58$\,ms to $9.30$\,ms, corresponding to an $89.2\times$ speedup. These results demonstrate an effective quality--efficiency trade-off for strict one-step text-to-motion generation.
CommentsUnder Review