arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.39575cs.ROcs.AI

ECHO-G: 具身共语拟人运动生成

ECHO-G: Embodied Co-speech Humanoid mOtion Generation

  • Beihang University(北京航空航天大学)
  • Mondo Robotics
  • The Hong Kong University of Science and Technology(香港科技大学)
  • Nanjing University(南京大学)

机构由 AI 辅助整理,请以论文原文为准。

Yizhao Li, Pusen Gao, Ming Wang, Shaojie Shen, Shuo Yang, Hao Xu

AI总结:

ECHO-G提出一种联合语音音频与定时文本的框架,通过扩散变压器直接在机器人空间生成共语全身运动,并配套数据集与基准,验证了优于重定向方法且可部署于实体机器人。

AI中文摘要:

生成拟人机器人的全身共语运动需要协调语音韵律、语言内容和具身特定运动。为此,我们提出了ECHO-G,一个以语音音频和定时文本为联合条件的框架。其语音接地扩散变压器(SGDiT)将帧对齐的声学特征与词元级语言上下文相结合,保留它们各自的粒度。通过整流流匹配训练,它直接在机器人空间中建模一对多的语句-运动关系。为支持训练和评估,我们引入了一个基于BEAT2的音频-文本-机器人数据集和一个涵盖共语特征、机器人运动质量和运行时效率的基准。对比评估支持直接在机器人空间中生成,优于所评估的人体运动生成和重定向流程,而模态消融研究突显了联合音频-文本条件的优势。我们进一步在实体拟人机器人上展示了部署。一项补充的视频评分研究也倾向于联合条件而非替代方案。数据集及训练、推理和评估代码可通过我们的项目页面获取。

英文摘要:

Generating full-body co-speech motion for humanoid robots requires coordinating speech prosody, linguistic content, and embodiment-specific motion. To this end, we present ECHO-G, a framework jointly conditioned on speech audio and timed transcripts. Its Speech-Grounded Diffusion Transformer (SGDiT) combines frame-aligned acoustic features with token-level linguistic context, preserving their distinct granularities. Trained with rectified flow matching, it models one-to-many utterance-motion relationships directly in robot space. To support training and evaluation, we introduce a BEAT2-derived audio-text-robot dataset and a benchmark covering co-speech characteristics, robot-motion quality, and runtime efficiency. Comparative evaluation supports direct robot-space generation over the evaluated human-motion generation and retargeting pipelines, while modality ablations highlight the benefits of joint audio-text conditioning. We further demonstrate deployment on a physical humanoid robot. A complementary video-rating study also favors joint conditioning over the alternatives. The dataset and training, inference, and evaluation code are available through our project page.

补充信息

↑