发表机构
The Chinese University of Hong Kong; Malmö University; National University of Sciences and Technology (NUST)(香港中文大学; 马尔默大学; 国立科技大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对可穿戴传感器HAR中多通道判别信息获取难、数据小易过拟合及现有增强方法领域依赖的问题,提出ALAE-TAE-CutMix+及扩展ALAE-CIE-TAE-CutMix+框架,增强潜在特征并引入新增强策略,在四个数据集上显著优于SOTA。
AI 中文摘要
基于可穿戴传感器的人类活动识别(HAR)通过其多种应用极大地改善了人类生活质量。对于HAR而言,多传感器通道信息对于获得最佳性能至关重要。当前研究表明,应用注意力神经网络来优先处理具有判别性的传感器通道有助于模型更精确地对活动进行分类。然而,从多传感器通道中获取判别性信息并非总是轻而易举的,例如在收集老年住院患者的数据时。在此背景下,现有的HAR方法难以对活动进行分类,尤其是性质相似的活动。此外,由于可用数据集规模较小,HAR模型主要面临过拟合问题,这导致性能不佳。数据增强(DA)是解决该问题的一种可行方案。然而,现有的DA方法存在各种缺陷,包括可能依赖特定领域,从而导致测试序列的模型失真。为解决这些HAR问题,我们提出了一种新颖框架ALAE-TAE-CutMix+,该框架侧重于两个方面。首先,它增强每个传感器通道的潜在信息,并学习利用多个潜在特征与当前活动之间的关系。因此,每个活动的判别性特征表示得到丰富。其次,引入一种新的增强策略以解决现有传感器通道数据增强的不足。随后,我们扩展该框架以创建进一步增强版本,即ALAE-CIE-TAE-CutMix+,该版本学习捕捉每对传感器通道特征之间的交互。我们发现,尽管第一个框架的性能略优于后者,但后者更为可靠和稳健。两个框架在来自不同领域的四个HAR数据集上均显著优于最先进(SOTA)方法。
英文摘要
Human Activity Recognition (HAR) through wearable sensors greatly improves the quality of human life through its multiple applications. For HAR, multi-sensor channel information is vital for optimal performance. Current work states that applying an attention neural network to prioritize discriminatory sensor channels helps the model classify activity more precisely. However, obtaining discriminatory information from multisensory channels is not always trivial, such as when collecting data from older hospitalized patients. In this context, existing HAR methods struggle to classify activities, particularly activities with similar natures. Moreover, HAR models predominantly suffer from overfitting due to the small size of available datasets, which leads to poor performance. Data augmentation (DA) is a viable solution to this problem. However, available DA methods have various drawbacks, including the possibility of being domain-dependent, resulting in distorted models for test sequences. To address these HAR problems, we propose a novel framework, ALAE-TAE-CutMix+, which focuses on two aspects. First, it enhances the latent information across each sensor channel and learns to exploit the relation among multiple latent features and the ongoing activity. Consequently, the discriminatory feature representations of each activity is enriched. Second, a new augmentation strategy is introduced to address the shortcomings of existing multi-sensor channel data augmentation. We then extend the framework to create a further enhanced version, namely ALAE-CIE-TAE-CutMix+, which learns to capture the interactions between the features of each pair of sensor channels. We find that although the first framework performs slightly better than the latter, the latter is nonetheless more reliable and robust. Both frameworks significantly outperform SOTA approaches on the four HAR datasets from diverse domains.