发表机构
Korea University; Purdue University(高丽大学; 普渡大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对第一人称视频中手部动态感知难的问题,提出基于功能手型先验的识别框架,通过手型分类与语义嵌入提升3D姿态估计和动作识别,在FPHA和H2O基准上超越现有方法。
AI 中文摘要
当前用于第一人称视角动作识别的方法,在仅依赖几何或物理信息时,往往难以感知动态手部运动。在本工作中,我们通过深入理解功能性手部构型与物体之间的关联,有效解决了这一问题,从而提升了对真实世界场景的细致解读能力。为此,我们基于功能视角引入了一种实用的手型分类体系,并将其用于对现有数据集进行逐帧手型标注。我们还提出了一种新颖的手部动作识别框架,将手型的语义细节作为先验信息。该方法增强了网络在整个动作序列中对连续手部交互的理解。我们的完整流程由三个主要模块组成:(1)特征提取模块,(2)第一人称知识模块,该模块利用短期线索估计3D手部姿态、物体类别和手型,以及(3)第一人称动作模块,该模块在更长时间范围内聚合逐帧知识,包括手型的文本嵌入。在我们使用大规模基准FPHA和H2O进行的广泛实验中,我们的模型优于当前最先进的方法,展示了其卓越的性能。
英文摘要
Current methods for egocentric view action recognition often face challenges in perceiving dynamic hand movements relying solely on geometrical or physical information. In this work, we effectively address this problem by gaining insights into the correlation between functional hand configurations and objects, which improves the detailed interpretation of real-world scenarios. To this end, we introduce a practical taxonomy of hand types based on the functioning perspective and utilize it for per-frame hand type labeling on existing datasets. We also propose a novel hand action recognition framework considering semantic details of the hand type as prior. This approach boosts the network's understanding of the continuous hand interaction throughout the action sequence. Our whole pipeline consists of three main modules: (1) Feature Extraction, (2) Egocentric Knowledge Module, which estimates 3D hand pose, object category, and hand type leveraging short-term cues, and (2) Egocentric Action Module, which aggregates per-frame knowledge, including text embeddings of hand type, over a longer time. In our extensive experiments with large-scale benchmarks, FPHA and H2O, our model outperforms current state-of-the-art methods, demonstrating its superior performance.
CommentsBMVC 2023 Oral Paper