AI 中文总结
针对近距离人机交互中类似躯干机器人关联开放词汇响应难的问题,提出基于轻量级流匹配的全身语义到驱动基础框架,将多模态大语言模型响应转换为元组,参数化轨迹并采样运动,实验显示其提升关联正确性、减少推理时间,还提升了人机交互满意度。
AI 中文摘要
对于近距离人机交互(HRI),类似躯干的连续体操纵器为多样化的全身表达提供了物理通道,但将开放词汇响应与此类机器人进行关联具有挑战性:末端执行器运动无法充分说明身体形状,而直接的全身命令是高维的且难以保持可行性。我们提出了一个基于轻量级流匹配的大象启发式软躯干HRI的全身语义到驱动的基础框架。该框架将多模态大语言模型的响应转换为有界的、形态对齐的意图强度元组,使用紧凑的Catmull-Rom样条控制对肌腱驱动轨迹进行参数化,并使用整流流生成器对可行的全身躯干运动进行采样。实验表明,所提出的框架将保留集的关联正确性从25.0%提高到77.2%,优于原始响应密集回归基线。与去噪扩散基线相比,它将正确性从71.9%提高到77.2%,并将推理时间从7.86毫秒减少到4.87毫秒,同时保留运动多样性。一项有100名参与者的物理HRI研究进一步表明,添加生成的软躯干运动通道将总体满意度从仅视听基线的46%提高到82%。
英文摘要
For close-contact human-robot interaction (HRI), trunk-like continuum manipulators provide a physical channel for diverse whole-body expression, but grounding open-vocabulary responses into such robots is difficult: end-effector motion underspecifies body shape, whereas direct whole-body commands are high-dimensional and hard to keep feasible. We propose a whole-body semantic-to-actuation grounding framework for elephant-inspired soft-trunk HRI based on lightweight flow matching. The framework converts responses from a multimodal large language model into bounded, morphology-aligned intent-intensity tuples, parameterizes tendon-actuation trajectories with compact Catmull-Rom spline controls, and uses a rectified-flow generator to sample feasible whole-body trunk motions. Experiments show that the proposed framework improves held-out grounding correctness from 25.0% to 77.2% over a raw-response dense-regression baseline. Compared with a denoising-diffusion baseline, it improves correctness from 71.9% to 77.2% and reduces inference time from 7.86 ms to 4.87 ms while preserving motion diversity. A 100-participant physical HRI study further shows that adding the generated soft-trunk motion channel increases the positive overall-satisfaction rating from 46% to 82% over the audiovisual-only baseline.