发表机构
Technological University of Uruguay; Ostfalia University of Applied Sciences(乌拉圭科技大学; 奥斯特法利亚应用科技大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出结合几何与语义感知层的混合注意力估计流程,通过有限状态机调节自适应人机交互,经10人40次试验验证其交互启动可靠、暂停行为一致且输出信息非冗余。
AI 中文摘要
本文针对基于InMoov生态系统的富有表现力的机器人头部的人机交互,开展了混合视觉注意力估计的应用案例研究。所提出的流程将快速几何感知层与基于视觉语言模型的独立语义感知层相结合:几何层提供高频的人脸与头部姿态信息用于时间调节;语义层仅接收原始自我中心相机帧,生成与机器人注意力、手机使用或其他注意力相关的上下文注意力标签。这些信号通过有限状态机整合,调节自适应交互行为,包括激活、等待、交互恢复和返回休息状态。该系统在10名参与者的40次试验中接受评估,涵盖基线和自适应交互条件。结果显示,所有试验中交互启动可靠,自适应干扰条件下暂停行为一致,且几何与语义输出间存在非冗余语义信息。
英文摘要
This paper presents an applied case study on hybrid visual attention estimation for human-robot interaction using an expressive robotic head based on the InMoov ecosystem. The proposed pipeline combines a fast geometric perception layer with an independent semantic perception layer based on a vision-language model. The geometric layer provides high-frequency face and head-pose information for temporal regulation, while the semantic layer receives only raw egocentric camera frames and produces contextual attention labels related to attention toward the robot, phone use, or attention elsewhere. These signals are integrated through a finite state machine that regulates adaptive interaction behavior, including activation, waiting, interaction resumption, and return to rest. The system was evaluated with 10 participants across 40 trials covering baseline and adaptive interaction conditions. Results show reliable interaction start across all trials, consistent pause behavior in the adaptive distraction condition, and non-redundant semantic information between the geometric and semantic outputs.