发表机构
Sapienza University of Rome; International University of Rome(罗马第一大学; 罗马国际大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对人形机器人自主性和动态环境响应受限问题,提出语义音频驱动的多模态编排框架,通过处理音频流、结合音乐与语音输入及强化学习控制,验证了该方法在模拟和实体机器人上的有效性。
AI 中文摘要
近期人形机器人技术和强化学习的进展使得获取高表现力的全身运动策略成为可能。然而,大多数机器人性能仍基于预编程序列或外部触发行为,限制了自主性和对动态环境的响应能力。本文引入了一种新颖的多模态编排框架,用于语义音频驱动的人形机器人控制,使机器人能够实时自主选择并执行合适的运动技能。该系统处理连续音频流并将其路由到音乐或语音分支。音乐输入通过音频指纹识别和语义嵌入来检索曲目身份和时间对齐,实现音乐片段与运动策略之间的动态映射。语音输入基于模仿学习技能的离散库,实现直接的人机交互。两种模态共享一个统一接口,通过强化学习控制管道调度技能执行。我们在模拟环境和Unitree G1人形机器人上验证了该方法,展示了强大的模拟到现实的迁移能力以及一致的音频条件策略选择。
英文摘要
Recent advances in humanoid robotics and reinforcement learning have enabled the acquisition of highly expressive whole-body motion policies. However, most robotic performances remain based on pre-scripted sequences or externally triggered behaviors, limiting autonomy and responsiveness to dynamic environments. In this work, we introduce a novel multi-modal orchestration framework for semantic audio-driven humanoid control, enabling robots to autonomously select and execute appropriate motion skills in real time. The system processes continuous audio streams and routes them into music or speech branches. Music input is handled via audio fingerprinting and semantic embeddings to retrieve track identity and temporal alignment, allowing dynamic mapping between musical segments and motion policies. Speech input is grounded into a discrete library of imitation-learned skills, enabling direct human-robot interaction. Both modalities share a unified interface that schedules skill execution over a reinforcement learning control pipeline. We validate the approach in simulation and on a Unitree G1 humanoid, showing robust sim-to-real transfer and consistent audio-conditioned policy selection. Supplementary materials are available at the following site: https://lab-rococo-sapienza.github.io/semantic-WBC/
CommentsAccepted at 29th Robocup International Symposium, held on July 6th, 2026 in Incheon, Republic of Korea