发表机构
State Key Laboratory of General Artificial Intelligence, Peking University, Shenzhen Graduate School; University at Buffalo SUNY; Department of Computer and Information Science, University of Pennsylvania(北京大学深圳研究生院通用人工智能国家重点实验室; 纽约州立大学布法罗分校; 宾夕法尼亚大学计算机与信息科学系)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对人机和人与人交互识别,提出STAR方法,通过骨骼令牌对齐与重排,利用实体重排、交互式时空令牌及视觉交互编码等组件,结合骨骼与RGB视频数据学习特征,经对比学习对齐表示,实验验证其性能优于现有方法。
AI 中文摘要
理解人机和人与人之间的物理交互是3D视觉中一个具有挑战性但正在兴起的主题。大多数现有方法依赖骨骼序列,虽在低光和隐私敏感环境中有效,但面临两大挑战:从骨骼数据学习并有效利用交互线索,以及弥补仅骨骼中缺乏的视觉信息。为应对这些挑战,我们提出用于人机和人与人交互识别的骨骼令牌对齐与重排(STAR)。它学习特定于交互的骨骼特征,并通过在共享潜在空间中对齐骨骼和RGB视频表示,利用视觉线索丰富这些特征。具体而言,STAR由三个关键组件组成。首先,设计一个使用实体重排(ER)和交互式时空令牌(IST)捕获细粒度相互依赖关系的骨骼编码器。其次,提出视觉交互编码,引入关注交互(FoI)策略以关注RGB视频中与交互相关的时空区域。最后,通过对比学习目标对齐这些表示,并使用细化头进一步细化预测。在训练期间,STAR利用骨骼和RGB视频数据学习强大的、有区分力的交互表示。在推理时,它仅在骨骼上运行,保留视觉信息带来的好处,同时保持仅骨骼的效率。在Chico、HARPER、NTU Mutual 11和26数据集上的大量实验通过展示优于现有方法的性能,一致验证了我们的方法。我们的代码可在这个https URL上公开获取。
英文摘要
Understanding physical human-robot and human-human interactions is a challenging yet emerging topic in 3D vision. While most existing methods rely on skeleton sequences--effective in low-light and privacy-sensitive environment--they face two major challenges: 1) learning and effectively exploiting interaction cues from skeletal data, and 2) compensating for the lack of visual information absent in skeletons alone. To address these challenges, we propose skeletal token alignment and rearrangement (STAR) for human-robot and human-human interaction recognition. It learns interaction-specific skeleton features and enriches them using visual cues by aligning skeleton and RGB video representations in a shared latent space. Specifically, STAR consists of three key components. First, we design a skeleton encoder that captures fine-grained interdependencies using Entity Rearrangement (ER) and Interactive Spatiotemporal Tokens (ISTs). Second, we present Visual Interaction Encoding that introduces a Focus on Interactions (FoI) strategy to attend to spatiotemporal regions relevant to interactions in RGB videos. Finally, these representations are aligned via a contrastive learning objective, with a refinement head further refines predictions. During training, STAR leverages both skeleton and RGB video data to learn robust, discriminative interaction representations. At inference time, it operates on skeletons alone, retaining visual-informed benefits while preserving skeleton-only efficiency. Extensive experiments on Chico, HARPER, NTU Mutual 11 and 26 datasets consistently validate our approach by demonstrating superior performance over state-of-the-art methods. Our code is publicly available at https://github.com/Necolizer/STAR.
CommentsAccepted for publication in IEEE Transactions on Multimedia (IEEE TMM)
Journal refIEEE Transactions on Multimedia, vol. 28, pp. 4652-4665, 2026