arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.23414cs.RO

EmoPose:视觉语言模型引导的仿人机器人情感感知手势生成

EmoPose: Vision-Language Model Guided Emotion-Aware Gesture Generation for Humanoid Robots

  • The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))
  • The Shandong University(山东大学)
  • RoboScience

机构由 AI 辅助整理,请以论文原文为准。

Daojie Peng, Bingtao Wang, Fulong Ma, Wenjun Yue, Liang Zhang, Jun Ma

AI总结:

EmoPose提出VLM引导的仿人机器人情感手势生成框架,通过可执行语义接口和动作库实现表达性与可控性,在基准上显著优于直接提示方法。

AI中文摘要:

具有社交能力的仿人机器人必须通过手势和言语来传达情感与意图,然而开放式交互需要生成既富有表现力又能在特定机体上执行的动作。这要求在保留确定性、具身感知的机器人控制的同时,具备语义灵活性以适应情境化的社交意图。我们提出了EmoPose,一个由视觉语言模型(VLM)引导的框架,通过可执行的语义接口弥合了这一差距。给定语言、对话历史以及可选的视觉上下文,VLM选择包含交际类别、库变体、强度和语音锚点的有序手势计划。一个可扩展的机器人自有动作库定义了可用的表达词汇以及14自由度关节目标的来源。Pose Studio支持自动轨迹生成、MuJoCo预览以及新库条目与VLM引导的自动同步;确定性的机器人侧模块验证计划、构建轨迹、调度手势并管理排队和中断。这种划分使得交互库能够针对新的社交情境进行扩展,而无需改变控制接口或将原始关节命令委托给基础模型。在EmoPose-Bench上,结构化的GPT-5.5规划在Easy层级达到$98.25\pm0.52\\%$,总体达到$76.50\pm0.54\\%$,超过了同模型的直接标签提示。进一步的测试验证了对话上下文使用和有序多动作组合。该系统完成了标称的MuJoCo测试套件,并在物理Unitree G1上实现了全部29个作者设计的变体。一次四站实验室参观展示了带有中断的表达性叙述、基于摄像头的对话以及导航。

英文摘要:

Socially competent humanoid robots must communicate affect and intent through gesture as well as speech, yet open-ended interaction must become motion that is both expressive and executable on a specific body. This demands semantic flexibility for contextual social intent while preserving deterministic, embodiment-aware robot control. We present EmoPose, a vision-language model (VLM)-guided framework that bridges this gap through an executable semantic interface. Given language, dialogue history, and optional visual context, the VLM selects an ordered gesture plan containing a communicative class, library variant, intensity, and speech anchor. A scalable robot-owned motion library defines the available expressive vocabulary and the source of 14-DoF joint targets. Pose Studio supports automatic trajectory generation, MuJoCo preview, and automatic synchronization of new library entries with the VLM guide; deterministic robot-side modules validate plans, construct trajectories, schedule gestures, and manage queueing and interruption. This division lets the interaction repertoire grow for new social contexts without changing the control interface or delegating raw joint commands to the foundation model. On the EmoPose-Bench, structured GPT-5.5 planning reaches $98.25\pm0.52\%$ on the Easy tier and $76.50\pm0.54\%$ overall, exceeding same-model direct-label prompting. Further tests validate dialogue-context use and ordered multi-action composition. The system completes the nominal MuJoCo suite and realizes all 29 authored variants on the physical Unitree G1. A four-stop laboratory tour demonstrates expressive narration with interruption, camera-grounded dialogue, and navigation.

↑