MASQ:面向无监督骨骼动作分割的掩码感知时空量化
MASQ: Mask-Aware Spatiotemporal Quantization for Unsupervised Skeleton Action Segmentation
- University of Science and Technology of China(中国科学技术大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
该研究针对无监督骨骼动作分割中空间掩码与离散量化引发的不稳定问题,提出MASQ框架,通过JLSD机制与掩码感知速度损失优化,在三个数据集上实现优于现有方法的帧均值准确率。
AI中文摘要:
无监督骨骼基时间动作分割是理解未修剪长序列中人类行为的关键任务。近期方法常依赖离散量化从运动表示中发现动作边界,但当引入空间掩码进行表示学习时,会产生表示歧义,而离散量化会进一步放大潜在空间中的微小波动,这两个因素的相互作用常导致代码切换不稳定及动作边界附近出现严重时间抖动。为解决这些限制,我们提出了一种新颖的掩码感知动作时空量化(MASQ)框架,该框架将空间特征推理与时间建模这两个冲突任务解耦。在空间维度,我们引入关节级结构化丢弃(JLSD)机制,对选定关节的整个时间轨迹进行掩码,以鼓励模型学习具有区分性的关节间协调模式;在时间维度,我们设计了掩码感知速度损失,仅对可见关节强制运动一致性,从而防止掩码信号导致的梯度冲突并稳定时间预测。在HuGaDB、LARa和BABEL三个广泛使用的骨骼数据集上进行的大量实验表明,所提出的MASQ框架显著优于现有的无监督最先进方法,尤其在帧均值准确率上建立了全面且实质性的领先优势。
英文摘要:
Unsupervised skeleton-based temporal action segmentation is a crucial task for understanding human behavior in long untrimmed sequences. Recent approaches often rely on discrete quantization to discover action boundaries from motion representations. However, when spatial masking is introduced for representation learning, it can introduce representation ambiguity, while discrete quantization further amplifies small fluctuations in the latent space. The interaction between these two factors often leads to unstable code switching and severe temporal jitter near action boundaries.To address these limitations, we propose a novel Mask-aware Action Spatiotemporal Quantization (MASQ) framework. Our framework decouples the conflicting tasks of spatial feature inference and temporal smoothing.In the spatial dimension, we introduce a Joint-Level Structured Dropout (JLSD) mechanism that masks the entire temporal trajectory of selected joints, to encourage the model to learn discriminative inter-joint coordination patterns. In the temporal dimension, we design a mask-aware velocity loss that enforces motion consistency only on visible joints, that prevents gradient conflicts caused by masked signals and stabilizing temporal predictions. Extensive experiments on three widely used skeleton datasets, including HuGaDB, LARa, and BABEL, demonstrate that the proposed MASQ framework significantly outperforms existing state-of-the-art unsupervised methods. In particular, our model establishes a comprehensive and substantial leading advantage in the Mean over Frames accuracy.