发表机构
Edge Hill University(埃奇希尔大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出FineX模型,融合RGB、姿态热图与骨骼图线索,通过交叉注意力和稀疏混合专家实现细粒度动作识别,在多数据集上取得最优,长尾Gym288准确率提升7.6个百分点。
AI 中文摘要
细粒度人体动作识别(FHAR)需区分视觉相似但在身体姿态、时序或局部外观上存在差异的动作。RGB表示保留视觉上下文但常抑制关节级几何信息,而骨骼表示编码运动学信息却丢失密集空间细节。我们提出FineX,将细粒度线索分解为RGB外观、姿态热图几何和骨骼图拓扑。成对交叉注意力实现对称的、流保持的信息交换,随后是流向潜在稀疏混合专家,将各表示路由至内容相关的共享专家子集,由负载均衡目标进行正则化。FineX在Gym99、Gym288和Diving48上取得了最优结果。在长尾Gym288上,它将平均类别准确率从68.6%提升至76.2%(+7.6个百分点),且无需文本监督或大规模视觉-语言预训练,证明了结构化视觉-姿态-图融合和条件专家优化对FHAR的益处。
英文摘要
Fine-grained human action recognition (FHAR) must distinguish visually similar actions that differ mainly in body configuration, timing, or local appearance. RGB representations retain visual context but often suppress joint-level geometry, whereas skeleton representations encode kinematics but discard dense spatial detail. We introduce FineX, which factorizes fine-grained cues into RGB appearance, pose heatmap geometry, and skeletal-graph topology. Pairwise cross-attention enables symmetric, stream-preserving information exchange, followed by a streamwise latent sparse Mixture-of-Experts that routes each representation to a content-dependent subset of shared experts, regularized by a load-balancing objective. FineX achieves state-of-the-art results on Gym99, Gym288, and Diving48. On the long-tailed Gym288, it raises mean class accuracy from 68.6% to 76.2% (+7.6 points) without textual supervision or large-scale vision-language pre-training, demonstrating the benefit of structured visual-pose-graph fusion and conditional expert refinement for FHAR.
CommentsAccepted manuscript: Workshop on Affective & Behavior Analysis in-the-wild (ABAW), as part of European Conference on Computer Vision (ECCV) 2026