用于人机协作中低延迟人类动作预测的姿态锚定光流
Pose-Anchored Optical Flow for Low-Latency Human Action Anticipation in Human-Robot Teaming
浏览论文内容
中文总结 AI 辅助
本文提出PoseOFF姿态锚定光流表示,用于人机协作中低延迟人类动作预测,可在早期观测下提升动作识别准确率,无需全帧运动处理,适用于实时资源受限场景。
中文摘要 AI 辅助
人机交互(HRI)要求机器人在人类动作执行初期就对其进行解读,以实现安全、高效且自然的响应。然而,许多现有的人类动作识别方法要么依赖稀疏的骨骼表示(缺乏细粒度运动线索),要么依赖密集光流(对于低延迟感知流水线而言计算成本高昂)。本文提出PoseOFF,一种姿态锚定光流表示,用于捕捉人类关节周围的局部运动信息,以支持更早的人类意图理解。通过将运动特征提取建立在人类姿态的条件下,PoseOFF在语义有意义的身体位置编码局部运动动态,形成与人类运动学明确对齐的结构化运动表示。我们在多个基准数据集和动作预测的骨干网络架构上对PoseOFF进行评估,证明其在识别准确率上的持续提升,尤其是在早期观测比例下。结果表明,PoseOFF能让模型在观测更少动作序列的同时实现相当或更优的性能,凸显其早期预测的有效性。重要的是,这些增益无需处理全帧运动即可实现,使该方法适用于实时和资源受限的场景。这些发现表明,诸如PoseOFF之类的以姿态为中心的运动表示可增强交互机器人系统更早推断人类动作的能力,为人机交互场景中更具响应性和预测性的行为提供支持。
英文摘要
Human-robot interaction (HRI) requires robots to interpret human actions early in their execution in order to respond safely, efficiently, and naturally. However, many existing approaches to human action recognition rely either on sparse skeletal representations, which lack fine-grained motion cues, or dense optical flow, which can be computationally expensive for low-latency perception pipelines. In this paper, we propose PoseOFF, a pose-anchored optical flow representation that captures local motion information around human joints to support earlier human intent understanding. By conditioning motion feature extraction on human pose, PoseOFF encodes localised motion dynamics at semantically meaningful body locations, forming a structured motion representation that is explicitly aligned with human kinematics. We evaluate PoseOFF across multiple benchmark datasets and backbone architectures for action anticipation, demonstrating consistent improvements in recognition accuracy, particularly at early observation ratios. Our results show that PoseOFF enables models to achieve comparable or improved performance while observing less of the action sequence, highlighting its effectiveness for early prediction. Importantly, these gains are achieved without requiring full-frame motion processing, making the approach practical for real-time and resource-constrained settings. These findings suggest that pose-centred motion representations such as PoseOFF can enhance the ability of interactive robot systems to infer human actions earlier, supporting more responsive and anticipatory behaviour in human-robot interaction scenarios.
发表机构
- Swinburne University of Technology(斯威本科技大学)
- School of Science, Computing and Emerging Technologies(科学、计算与新兴技术学院)
机构由 AI 辅助整理,请以论文原文为准。