DiMoP:基于帧级伪分类的扩散驱动运动表示学习用于基于骨架的动作识别
DiMoP: Diffusion-Driven Motion Representation Learning With Frame-Level Pseudo-Classification for Skeleton-Based Action Recognition
浏览论文内容
中文总结 AI 辅助
DiMoP提出掩蔽与扩散驱动的运动表示学习方法,结合帧级伪分类,均匀学习弱中强运动,在NTU RGB+D 60/120和PKUMMD上达到最先进性能。
中文摘要 AI 辅助
鲁棒的基于骨架的动作识别需要能够捕捉从细微、适中到强烈运动的广泛光谱的表示。现有方法往往侧重于强烈运动。本文介绍了DiMoP,一种掩蔽和扩散驱动的运动表示学习方法,采用帧级伪分类,显式学习关节运动的分布,而非像现有方法那样回归确定性坐标。通过用渐进噪声扩散掩蔽的关节,并在可见关节条件下对其进行去噪,DiMoP通过可控的加噪和去噪过程进行学习,实现对弱、中和强动态的均匀学习。为了赋予基于掩蔽的生成式扩散学习以判别能力,提出了一个伪帧分类器,强制学习朝向序列一致和时间连贯的伪标签,无需人工标注。这些策略共同为生成式和判别式运动建模提供了原则性机制。DiMoP在NTU RGB+D 60/120和PKUMMD上取得了最先进的性能,包括在NTU RGB+D 120的跨受试者协议上比先前工作提升了1.1个百分点。
英文摘要
Robust skeleton-based action recognition requires representations that capture a wide spectrum of motions, from subtle to moderate and strong ones. Existing methods often focus on strong motions. This paper introduces DiMoP, a masking- and diffusion-driven motion representation learning method with frame-level pseudo-classification to explicitly learn the distribution of joint motions rather than regressing deterministic coordinates, as existing methods often do. By diffusing masked joints with progressive noise and denoising them conditioned on visible joints, DiMoP learns through controllable noising and denoising processes, enabling uniform learning of weak, moderate, and strong dynamics. To enable the masking-based generative diffusion learning with a discriminative capability, a pseudo-frame classifier is proposed that enforces the learning towards sequence-consistent and temporally coherent pseudo-labels without manual annotations. Together, these strategies provide a principled mechanism for joint generative and discriminative motion modeling. DiMoP achieves state-of-the-art performance across NTU RGB+D 60/120, and PKUMMD, including a 1.1 percentage point gain over prior works on NTU RGB+D 120 with the cross-subject protocol.
发表机构
- University of Wollongong(伍伦贡大学)
机构由 AI 辅助整理,请以论文原文为准。