SocialHumanoid:通过一步式共语动作生成实现富有表现力的人形机器人行为
SocialHumanoid: Towards Expressive Humanoid Behavior via One-Step Co-Speech Motion Generation
浏览论文内容
中文总结 AI 辅助
SocialHumanoid提出一步式共语动作生成系统,实现情感可控、低延迟的人形机器人全身动作生成,并引入AffectMoCap数据集提升情感表现力,在BEAT2上取得最优FGD且推理速度约为GestureLSM的6倍。
中文摘要 AI 辅助
人形机器人日益被期望作为具身社交代理,通过面对面互动与人类自然交流。在此类交流中,人形机器人需要与语音同步、具有情感表现力且适合实时执行的身体行为。然而,现有的共语动作方法主要针对数字人开发,缺乏对情感控制和物理实体上低延迟连续生成的联合支持。为弥合这一差距,我们提出了SocialHumanoid,一个通过一步式共语动作生成实现富有表现力的人形机器人行为的系统。给定响应语音和指定的情感条件,SocialHumanoid在单次前向传播中生成每个全身动作窗口,并通过动作历史条件连接连续窗口。生成的人体动作进一步在线转换为适合实体机器人的参考轨迹,并由全身控制器跟踪以进行物理执行。为提供情感身体表达的显式监督,我们进一步引入了AffectMoCap,一个从两位专业演员处捕捉的4小时数据集,包含同步语音、身体动作、精细手部动作和情感标注。在BEAT2上,SocialHumanoid在比较的生成方法中取得了最佳FGD,具有竞争力的语音-动作同步性,并且在相同协议下推理速度约为GestureLSM的6倍。感知评估进一步表明,使用AffectMoCap训练提高了从生成身体动作中识别情感的能力,而真实机器人实验展示了连续情感条件行为和稳定的长时程执行。我们的项目页面见此URL。
英文摘要
Humanoid robots are increasingly expected to serve as embodied social agents that communicate naturally with humans through face-to-face interaction. During such communication, humanoid robots require body behaviors that are synchronized with speech, affectively expressive, and suitable for real-time execution. However, existing co-speech methods are primarily developed for digital humans and lack joint support for affective control and low-latency continuous generation on physical embodiments. To bridge this gap, we present SocialHumanoid, a system for expressive humanoid behavior via one-step co-speech motion generation. Given response speech and a specified affective condition, SocialHumanoid generates each full-body motion window in a single forward pass and connects successive windows through motion-history conditioning. The generated human motion is further converted online into embodiment-compatible robot references and tracked by a whole-body controller for physical execution. To provide explicit supervision for affective body expression, we further introduce AffectMoCap, a 4-hour dataset captured from two professional actors, containing synchronized speech, body motion, fine-grained hand motion, and emotion annotations. On BEAT2, SocialHumanoid achieves the best FGD among the compared generation methods, competitive speech-motion synchrony, and approximately $6\times$ faster inference than GestureLSM under the same protocol. Perceptual evaluations further show that training with AffectMoCap improves affect recognition from generated body motion, while real-robot experiments demonstrate continuous affect-conditioned behavior and stable long-horizon execution. Our project page is https://rex0191.github.io/SocialHumanoid/.