ECHO-G: 具身共语拟人运动生成
ECHO-G: Embodied Co-speech Humanoid mOtion Generation
- Beihang University(北京航空航天大学)
- Mondo Robotics
- The Hong Kong University of Science and Technology(香港科技大学)
- Nanjing University(南京大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
ECHO-G提出一种联合语音音频与定时文本的框架,通过扩散变压器直接在机器人空间生成共语全身运动,并配套数据集与基准,验证了优于重定向方法且可部署于实体机器人。
AI中文摘要:
生成拟人机器人的全身共语运动需要协调语音韵律、语言内容和具身特定运动。为此,我们提出了ECHO-G,一个以语音音频和定时文本为联合条件的框架。其语音接地扩散变压器(SGDiT)将帧对齐的声学特征与词元级语言上下文相结合,保留它们各自的粒度。通过整流流匹配训练,它直接在机器人空间中建模一对多的语句-运动关系。为支持训练和评估,我们引入了一个基于BEAT2的音频-文本-机器人数据集和一个涵盖共语特征、机器人运动质量和运行时效率的基准。对比评估支持直接在机器人空间中生成,优于所评估的人体运动生成和重定向流程,而模态消融研究突显了联合音频-文本条件的优势。我们进一步在实体拟人机器人上展示了部署。一项补充的视频评分研究也倾向于联合条件而非替代方案。数据集及训练、推理和评估代码可通过我们的项目页面获取。
英文摘要:
Generating full-body co-speech motion for humanoid robots requires coordinating speech prosody, linguistic content, and embodiment-specific motion. To this end, we present ECHO-G, a framework jointly conditioned on speech audio and timed transcripts. Its Speech-Grounded Diffusion Transformer (SGDiT) combines frame-aligned acoustic features with token-level linguistic context, preserving their distinct granularities. Trained with rectified flow matching, it models one-to-many utterance-motion relationships directly in robot space. To support training and evaluation, we introduce a BEAT2-derived audio-text-robot dataset and a benchmark covering co-speech characteristics, robot-motion quality, and runtime efficiency. Comparative evaluation supports direct robot-space generation over the evaluated human-motion generation and retargeting pipelines, while modality ablations highlight the benefits of joint audio-text conditioning. We further demonstrate deployment on a physical humanoid robot. A complementary video-rating study also favors joint conditioning over the alternatives. The dataset and training, inference, and evaluation code are available through our project page.