RoboGesture:面向人形机器人交互的实时语义对齐式伴随语音手势生成
RoboGesture: Real-Time Semantic-aligned Co-Speech Gestures Generation for Humanoid Interaction
另 3 家 · 查看机构详情
- Tsinghua University(清华大学)
- Galbot Inc.(Galbot公司)
- Beijing Institute of Technology(北京理工大学)
- Harbin Institute of Technology(哈尔滨工业大学)
- Peking University(北京大学)
- Shanghai Qi Zhi Institute(上海智知研究院)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
本研究提出 RoboGesture 框架,构建专属数据集并设计相关算法,使人形机器人可实时生成语义对齐的伴随语音手势,在 Unitree G1 机器人上的实验表明其性能优于现有最优方法。
中文摘要 AI 辅助
使人形机器人能够以同步且语义有意义的手势响应人类语音,是实现自然人机交互的基础。然而,该任务面临三大关键障碍:语义丰富的数据集稀缺、模型因运动学惯性而忽略音频线索的“模态 eclipse”(模态 eclipse 指模态遮蔽现象),以及涉及物理安全性的“仿真到现实 gap”(仿真到现实 gap 指仿真到现实的鸿沟)。我们提出 RoboGesture,这是一种以机器人为中心的框架,通过协同设计数据、建模与控制,打造出一套完整的人机交互系统,使机器人能够实时完成倾听、响应与手势生成。我们首先构建了包含 300 余种手势类别的 RoboGesture 数据集,并开发了一套自动化流程,用于合成大规模无碰撞、机器人专属的音频-运动配对数据。我们的架构包含一个分层语义-声学对齐器,可直接从原始音频 token 中提取多粒度韵律与语义线索;这些线索驱动基于带条件流匹配的扩散 Transformer 的流式条件运动生成器。为确保高响应性,我们引入了抗惯性 CFG 掩码,通过迫使模型主动从音频模态中挖掘控制信号,防止其坍缩为重复的历史模式。最后,基于 MPC 的安全过滤器确保在物理硬件上实现实时无碰撞执行。在 Unitree G1 人形机器人上开展的实验表明,与现有最优基线相比,RoboGesture 生成的响应更安全、更具节奏感且语义更恰当。
英文摘要
Enabling humanoid robots to respond to human speech with synchronized and semantically meaningful gestures is fundamental to natural human-robot interaction. However, this task faces three critical barriers: the scarcity of semantically rich datasets, the "modality eclipse" where models ignore audio cues in favor of kinematic inertia, and the sim-to-real gap regarding physical safety. We propose RoboGesture, a robot-centric framework that co-designs data, modeling, and control to power a complete interactive human-humanoid system in which the robot listens, responds, and gestures in real time. We first establish the RoboGesture dataset featuring over 300 gesture categories and develop an automated pipeline to synthesize large-scale collision-free, robot-specific audio-motion pairs. Our architecture features a Hierarchical Semantic-Acoustic Aligner that extracts multi-granular prosodic and semantic cues directly from raw audio tokens. These cues drive a Streaming Conditional Motion Generator based on a diffusion transformer with conditional flow matching. To ensure high responsiveness, we introduce Anti-Inertia CFG Masking, which prevents the model from collapsing into repetitive historical patterns by compelling it to proactively mine control signals from the audio modality. Finally, an MPC-based safety filter ensures real-time, collision-free execution on physical hardware. Experiments on a Unitree G1 humanoid demonstrate that RoboGesture generates safer, more rhythmic, and more semantically appropriate responses compared to state-of-the-art baselines.