arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

HALO:一种用于开放集人体活动识别的异构性感知语言对齐IMU基础模型

HALO: A Heterogeneity-Aware Language-Aligned IMU Foundation Model for Open-Set Human Activity Recognition

Zihan Ding, Liyu Zhang, Xiaomin Ouyang

arXiv 2608.27233首次发表:更新:

发表机构

Hong Kong University of Science and Technology(香港科技大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出HALO异构性感知语言对齐IMU基础模型,通过两阶段训练解决HAR的异构性与泛化问题,在多数据集上优于基线,零样本准确率提升13.7个百分点。

AI 中文摘要

基于惯性测量单元(IMU)的人体活动识别(HAR)可实现广泛应用,但该领域仍缺乏能在不同主体、设备和活动间泛化的统一模型。训练此类模型面临两大关键挑战:传感异构性(采样率、通道配置、传感器位置的差异),以及对未见活动和标签词汇的泛化能力差。我们提出HALO(异构性感知语言对齐开放集模型),一种针对IMU的领域专用基础模型,通过两阶段训练框架解决上述挑战。第一阶段采用异构性感知自监督学习预训练IMU编码器,包括自适应池化分词、通道无关特征提取,以及将自然语言传感器描述注入各通道嵌入的上下文传感器条件化;第二阶段通过同义词感知软对比学习将该IMU编码器与文本嵌入对齐,无需针对各数据集的分类器即可通过余弦相似度检索实现开放集识别。HALO在10个公开HAR数据集上训练,在7个保留数据集上评估,在所有8个综合指标上均优于5个最新基线模型,且在基线匹配输入下的4种设置中有3种仍领先。尽管仅使用约3500万可训练参数(比最新基础模型MOMENT的3.412亿少10倍),HALO在所有87个训练标签上测量的零样本开放集准确率提升了13.7个百分点。在另外两个存在严重分布偏移的数据集上,包括HALO在内的所有模型的零样本性能均出现崩溃。HALO在现实世界中性能的视频演示可在该https URL获取。

英文摘要

Human Activity Recognition (HAR) using inertial measurement units (IMUs) enables a wide range of applications, yet the field still lacks a unified model that can generalize across diverse subjects, devices, and activities. Training such a model is difficult due to two key challenges: sensing heterogeneity -- differences in sampling rates, channel configurations, and sensor placements -- and poor generalization to unseen activities and label vocabularies. We introduce HALO (Heterogeneity-Aware Language-aligned Open-set model), a domain-specific IMU foundation model that addresses both challenges through a two-stage training framework. Stage 1 pretrains the IMU encoder with heterogeneity-aware self-supervised learning, including adaptive-pooling tokenization, channel-independent feature extraction, and contextualized sensor conditioning that injects natural-language sensor descriptions into each channel embedding. Stage 2 aligns this IMU encoder with text embeddings via synonym-aware soft contrastive learning, enabling open-set recognition via cosine-similarity retrieval without per-dataset classifiers. Trained on 10 public HAR datasets and evaluated on 7 held-out datasets, HALO outperforms five state-of-the-art baselines on all 8 aggregate metrics, and still leads on 3 of 4 settings under baseline-matched inputs. Despite using only ~35M trainable parameters -- 10x fewer than the latest foundation model MOMENT (341.2M) -- HALO improves zero-shot open-set accuracy, measured over all 87 training labels, by 13.7 percentage points. On two further datasets with severe distribution shift, every model including HALO collapses zero-shot. A video demonstration of HALO's performance in real world is available at https://youtu.be/rooVKragtFU

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑