发表机构
Peking University; The Chinese University of Hong Kong, Shenzhen(北京大学; 香港中文大学(深圳))
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
Motion-Omni是端到端联合语音与全身动作生成框架,基于Qwen2.5-7B-Instruct实现,在动作指标接近教师级联的同时速度提升5.4倍,词错误率为全模态系统最低,还发布了相关数据集与评估协议。
AI 中文摘要
能够进行对话的虚拟角色需同时决定发言内容与发言时的动作,但相关能力分属不同模型家族:口语对话模型仅生成语音不生成动作,而协同语音动作模型仅能基于输入音频生成动作。现有标准解决方案是级联架构,先生成语音响应,再基于生成的音频运行动作模型,该方案需第二次完整推理,且无法实现两者的联合优化。本文提出Motion-Omni,这是一种端到端框架,其中口语对话模型可直接从生成语音的隐藏状态输出明确的面部表情,以及手部、上半身和下半身动作。联合训练在此处是必要的:冻结语音通路时,动作会与音频失配,需同时优化大语言模型(LLM)、语音生成器和动作生成器的两个目标,才能恢复对齐并保留口语对话能力。监督信号来自可扩展、与模型无关的流水线,该流水线用可替换的动作教师为一致性语音响应生成伪标签,得到422856个质量排序的配对(共1402小时)。我们还发布了SwDA-500,以及据我们所知首个面向随机开放型全身口语对话的公开评估协议,该协议匹配不同动作系统的音频,同时统一了渲染、自动指标、人工评估和延迟测量。以Qwen2.5-7B-Instruct为骨干实例化的Motion-Omni-Q7,在无参考动作指标上与同音频教师级联的性能差距在2%以内,响应速度快5.4倍(实时因子RTF=0.78,快于实时),在节拍相关性和多样性上超过所有非教师级联方案,且词错误率为2.62%,是所有对比全模态系统中最低的。
英文摘要
An avatar that holds a conversation should decide what to say and to move while saying it, yet these abilities live in separate model families: spoken dialogue models produce speech without motion, and co-speech motion models produce motion only from audio handed to them. The standard remedy is a cascade that first generates the spoken response and then runs a motion model over the finished audio, which requires a second full inference pass and precludes any joint optimisation between the two. We present Motion-Omni, an end-to-end framework in which a spoken dialogue model natively outputs explicit facial expression together with hand, upper-body and lower-body motion, generated directly from the hidden states that produce the speech. Joint training is not optional here: with the speech pathway frozen, motion remains misaligned with the audio, and co-adapting the LLM, Speech Generator and Motion Generator under both objectives is what recovers alignment while retaining spoken-dialogue ability. Supervision comes from a scalable, model-agnostic pipeline that pseudo-labels consistent-voice speech responses with a replaceable motion teacher, yielding 422,856 quality-ranked pairs (1,402 hours). We further release SwDA-500 and, to our knowledge, the first public evaluation protocol for stochastic open-ended full-body spoken dialogue, matching audio across motion systems while unifying rendering, automatic metrics, human evaluation, and latency measurement. Instantiated with a Qwen2.5-7B-Instruct backbone, Motion-Omni-Q7 matches the same-audio teacher cascade to within 2% on reference-free motion metrics while responding 5.4 x faster (RTF=0.78, faster than real time), surpasses all non-teacher cascades on beat correlation and diversity, and reaches a 2.62% word error rate, the lowest among the omni-modal systems compared.
Comments30 pages, 6 figures, 12 tables. Updated figures, presentation, and author notes. Project page: https://step-out.github.io/Motion-Omni-Page/ Code: https://github.com/step-out/Motion-Omni Data: https://huggingface.co/datasets/ChengqianMa/Motion-Omni