VSMP-IMU:基于视频的语义运动程序,用于生成传感器感知的合成IMU数据
VSMP-IMU: Video-Grounded Semantic Motion Programs for Sensor-Aware Synthetic IMU Generation
浏览论文内容
中文总结 AI 辅助
本文提出VSMP-IMU框架,基于结构化SMP生成可控合成IMU数据,在五个公开IMU-HAR数据集的多场景评估中,显著提升了HAR模型的性能。
中文摘要 AI 辅助
可穿戴人体活动识别(HAR)常受限于标注传感器数据的稀缺性,尤其在低资源、类别不平衡及跨主体泛化场景中。合成IMU生成可降低对标注数据的依赖,提升HAR机器学习模型性能,但现有方法存在权衡问题且未解决所有因素:视频驱动方法视觉上有依据,但对姿态估计误差敏感;文本驱动方法可控,但对活动实际执行方式的依据较弱。本文提出VSMP-IMU,一种基于视频的可控合成IMU生成框架,其基于结构化语义运动程序(SMP),将活动定义语义与标签保留变异分离。给定输入视频,VSMP-IMU提取并增强SMP,利用其合成运动,将运动转换为虚拟IMU信号,并将生成的信号适配到目标可穿戴领域。我们在五个公开IMU-HAR数据集上采用留一主体评估,将VSMP-IMU与现有最先进的合成数据生成方法对比。VSMP-IMU的平均Macro-F1达78.33%,较仅用真实数据训练提升9.77%,较最强的现有合成基线提升4.04%。在训练样本减少的低资源场景中,其较仅用真实数据训练提升18.54%,较最强的现有合成基线平均提升超6%。在不平衡数据集的长尾评估中,其尾部类别Macro-F1较仅用真实数据训练提升19.86%,较SOTA提升4.76%。这些结果表明,结构化的视频依据语义为生成可控、与可穿戴相关的合成传感器数据提供了实用基础。
英文摘要
Wearable human activity recognition (HAR) is often limited by the scarcity of labeled sensor data, especially in low-resource, class-imbalanced, and subject-generalization settings. Synthetic IMU generation can reduce this dependency and enhance HAR machine learning model's performance, but existing approaches face a trade-off without addressing all factors: video-driven methods are visually grounded but sensitive to pose-estimation errors, while text-driven methods are controllable but often weakly grounded in how activities are actually performed. We present VSMP-IMU, a video-grounded framework for controllable synthetic IMU generation based on a structured Semantic Motion Program (SMP), which separates activity-defining semantics from label-preserving variation. Given an input video, VSMP-IMU extracts and augments an SMP, uses it to synthesize motion, converts the motion into virtual IMU signals, and grounds the resulting signals to the target wearable domain. We evaluate VSMP-IMU against state-of-the-art synthetic data generation methods on five public IMU-HAR datasets under leave-one-person-out evaluation. VSMP-IMU achieves an average Macro-F1 of 78.33%, improving over real-only training by 9.77% and over the strongest prior synthetic baseline by 4.04%. In low-resource settings with reduced training data-samples, it improves over real-only training by 18.54% and over the strongest prior synthetic baselines by more than 6% on average. Under long-tail evaluation in imbalanced datasets, it improves tail-class Macro-F1 by 19.86% over Real-only training and by 4.76% over SOTA. These results show that structured video-grounded semantics provide a practical foundation for controllable, wearable-relevant synthetic sensor data generation.
发表机构
- DFKI(德国人工智能研究中心)
- RPTU Kaiserslautern-Landau(凯撒斯劳滕-兰道工业大学)
机构由 AI 辅助整理,请以论文原文为准。