发表机构
Tsinghua University; Beihang University; Durham University; Zhongguancun Laboratory(清华大学; 北京航空航天大学; 杜伦大学; 中关村实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
NeuroPath受人类感知通路启发,采用双通路图卷积网络架构,通过分组图卷积块和通路间融合模块提升基于骨骼的动作识别性能,在三个公开数据集上取得一致效果提升。
AI 中文摘要
基于骨骼的动作识别旨在从人体关节坐标序列中识别人类动作。大多数现有的时空图卷积网络(STGCNs)通过用隐式时空表示建模骨骼结构,已取得了良好的结果。然而,我们的实证研究显示不同骨骼模态间存在明显的性能不平衡,表明隐式耦合时空信息限制了对互补结构和运动线索的充分利用。受人类感知中的腹侧和背侧通路启发,我们提出了双通路图卷积网络(NeuroPath),采用双通路架构对时空信息进行分离但协同的建模。具体而言,变换单元首先将输入转换为通路特定的骨骼表示,使每个通路能专注于人类运动的互补方面。为进一步捕捉协调的关节行为及其相互关系,我们引入了分组图卷积块,该块可动态识别关键身体部位并建模其时空依赖关系。此外,通路间动态融合模块整合了跨通路的互补模态间信息,促进对动作的高层语义解释。在Kinetics Skeleton 400、NTU RGB+D 60和NTU RGB+D 120上进行的大量实验显示出一致的性能提升,验证了用于基于骨骼的动作识别的双通路时空建模的有效性。
英文摘要
Skeleton-based action recognition aims to recognize human actions from sequences of human joint coordinates. Most existing Spatial-Temporal Graph Convolutional Networks (STGCNs) have achieved promising results by modeling skeletal structures with implicit spatial-temporal representations. However, our empirical study reveals a clear performance imbalance across different skeletal modalities, indicating that implicitly coupling spatial and temporal information limits the full exploitation of complementary structural and motion cues. Inspired by the ventral and dorsal pathways in human perception, we propose Dual-Pathway Graph Convolutional Networks (NeuroPath), which adopt a dual-pathway architecture for separate yet collaborative modeling of spatial and temporal information. Specifically, transformation units first convert the input into pathway-specific skeletal representations, allowing each pathway to focus on complementary aspects of human motion. To further capture coordinated joint behaviors and their interrelationships, we introduce a group graph convolution block that dynamically identifies key body parts and models their spatial-temporal dependencies. In addition, inter-pathway dynamic fusion modules integrate complementary inter-modal information across pathways, facilitating higher-level semantic interpretation of actions. Extensive experiments on Kinetics Skeleton 400, NTU RGB+D 60, and NTU RGB+D 120 demonstrate consistent performance improvements, validating the effectiveness of dual-pathway spatial-temporal modeling for skeleton-based action recognition.
CommentsAccepted to Pattern Recognition
Journal refPattern Recognition, 2026
DOI:10.1016/j.patcog.2026.114723