发表机构
School of Computer Sciences, Universiti Sains Malaysia; Al-Quds Open University(计算机科学学院,马来西亚科学大学; 阿卡德大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对基于传感器的人类活动识别依赖大量标注数据的问题,提出联合嵌入预测架构框架,通过改进的VICReg目标函数从未标注数据学习通用表示,在两个基准数据集评估中成功学习高质量表示,在少数高方差过渡活动上泛化能力强。
AI 中文摘要
基于传感器的人类活动识别(HAR)在全监督学习环境中取得了显著进展。然而,这些监督学习模型依赖大量标注数据,需要大量人力收集和精细标注。为应对这些挑战,本文提出了一种专为基于传感器的HAR量身定制的联合嵌入预测架构框架,旨在从未标注数据集中学习强大且通用的表示。该框架有一个编码器,可明确对单个窗口内的细粒度局部时间表示以及相邻窗口的长期时间序列进行建模。此外,引入了改进的方差 - 不变性 - 协方差正则化(VICReg)目标函数,包含计算轻量级范数项以稳定JEPA预训练阶段。此函数平衡方差、不变性和协方差约束以防止表示崩溃。使用两个基准连续执行活动数据集对所提出的HAR - JEPA框架进行评估。结果表明该框架成功学习到了高质量表示。此外,HAR - JEPA学习到的表示在少数高方差过渡活动(如坐立和坐卧)上展现出卓越的泛化能力,而监督学习在这些活动上因支持有限容易过拟合。
英文摘要
Sensor-based human activity recognition (HAR) has achieved significant progressed in fully supervised learning settings. However, these supervised learning models rely on large amount of labeled data, which require labor-intensive collection and meticulous annotation. To address these challenges, this paper proposes a Joint Embedding Predictive Architecture framework tailored for sensor-based HAR, designed to learn robust and generalizable representations from unlabeled datasets. The proposed framework features an encoder designed to explicitly model both the fine-grained local temporal representations within individual window and the long-term temporal sequence of adjacent windows. Furthermore, we introduce an improved Variance-Invariance-Covariance Regularization (VICReg) objective function that incorporates computationally lightweight norm term to stabilize the JEPA pre-training phase. This term balances variance, invariance and covariance constraints to prevent representation collapse. The proposed HAR-JEPA framework is evaluated using two benchmark continuously performed activity datasets. The results show that high-quality representations are successfully learned by the proposed framework. Furthermore, the representations learned by HAR-JEPA demonstrates superior generalization on minority, high variance transitional activities such as sit-to-stand and sit-to-lie where supervised learning tend to overfit due to limited support.