发表机构
United Arab Emirates University (UAEU)(阿联酋大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对人形机器人交互中语音、面部和手势不协调的问题,提出TRABot框架,通过语义运动原子和组合规划器实现同步,实验证明其在自然性、表现力和多模态连贯性上最优。
AI 中文摘要
富有表现力的人形交互需要语音、面部动画和身体手势形成连贯的响应。然而,许多全身人形机器人产生语音和手势时没有视觉表现力的面部,而说话面部动画和机器人手势生成通常是分开开发的。我们提出了Talk, Render, Act(TRABot),一个基于智能体的框架,包含用于运动原子构建、对话生成、运动规划和面部动画的专门智能体。首先,为了产生自然且语义上有意义的手势,我们通过将长形式、G1重定向的BEAT2运动分割成具有自然手势边界、人工验证的交际功能和可行轨迹的单元,构建了机器人就绪的语义运动原子。其次,为了保持语义顺序并将身体运动与口语响应协调,我们引入了一个语义条件组合规划器。给定有序的语义功能序列和估计的响应持续时间,规划器选择已批准的原子以实现最长的可行动作序列,同时考虑过渡和中性恢复。最后,我们在物理G1人形机器人上部署了流式面部-语音-身体集成系统,在统一的实时交互循环中结合流式对话音频、音频驱动的面部动画和语义规划的身体运动。定量和定性实验表明,TRABot在自然性、表现力和多模态连贯性方面在所有比较条件下实现了最佳整体性能。
英文摘要
Expressive humanoid interaction requires speech, facial animation, and body gestures to form a coherent response. However, many full-body humanoid robots produce speech and gestures without a visually expressive face, while talking-face animation and robot gesture generation are typically developed separately. We present Talk, Render, Act (TRABot), an agent-based framework comprising specialized agents for motion-atom construction, dialogue generation, motion planning, and facial animation. First, to produce natural and semantically meaningful gestures, we construct Robot-Ready Semantic Motion Atoms by segmenting long-form, G1-retargeted BEAT2 motion into units with natural gesture boundaries, human-verified communicative functions, and feasible trajectories. Second, to preserve semantic order and coordinate body motion with the spoken response, we introduce a Semantic-Conditioned Compositional Planner. Given an ordered semantic function sequence and an estimated response duration, the planner selects approved atoms to realize the longest feasible action sequence while accounting for transitions and neutral recovery. Finally, we deploy a Streaming Face-Speech-Body Integration system on a physical G1 humanoid, combining streaming dialogue audio, audio-driven facial animation, and semantically planned body motion in a unified real-time interaction loop. Quantitative and qualitative experiments demonstrate that TRAbot achieves the best overall performance among all compared conditions in terms of naturalness, expressiveness, and multimodal coherence.