基于签名的特征学习用于人类活动识别:一项关于表示、深度和模型选择的可复现机器学习研究
Signature-Based Feature Learning for Human Activity Recognition: A Reproducible Machine Learning Study of Representation, Depth, and Model Choice
浏览论文内容
中文总结 AI 辅助
本研究在UCI HAR数据集上通过可复现且防泄漏的流程,系统评估了基于签名的特征学习对活动识别的影响,发现与随机森林结合时性能最优,最佳配置达到0.858准确率,且收益依赖于表示与分类器的匹配。
中文摘要 AI 辅助
人类活动识别(HAR)依赖于将传感器信号转换为信息丰富的表示以进行分类。尽管深度学习和手工特征被广泛使用,但表示本身的作用往往未被系统地单独考察。签名变换提供了一种数学上严谨的方法来编码时间顺序和跨通道交互,但其在完全可复现且防泄漏框架下对HAR的价值仍不明确。为评估基于签名的特征学习相较于原始信号基线是否能提升HAR性能,并考察嵌入策略、截断深度、模型选择和传感器配置的影响,我们在UCI HAR数据集上进行了实验,采用完全可复现的流程,保留原始训练-测试划分,并使用受试者不相交的验证以防止泄漏。我们比较了三种表示:原始展平信号、时间增强路径和超前-滞后变换路径。在多个截断深度下计算签名特征,并在相同的预处理和验证流程下使用多层感知机(MLP)和随机森林(RF)分类器进行评估。此外,还复现了一项先前基于K均值特征缩减的研究以作比较。基于签名的表示在与RF模型配合时提升了性能,最佳配置为使用基于熵的RF在深度6下的时间增强六通道签名,达到0.858的准确率和0.859的宏F1分数,优于最强的原始基线(准确率0.816)。超前-滞后表示在中等深度下具有竞争力,但未超越最佳的时间增强模型。MLP模型未超过原始基线。基于签名的特征学习可以改善HAR,但其收益取决于表示设计与分类器选择之间的对齐。
英文摘要
Human activity recognition (HAR) relies on transforming sensor signals into informative representations for classification. Although deep learning and handcrafted features are widely used, the role of representation itself is often not systematically isolated. Signature transforms provide a mathematically grounded way to encode temporal order and cross-channel interactions, but their value for HAR under a fully reproducible and leakage-aware framework remains unclear. To evaluate whether signature-based feature learning improves HAR performance compared with raw-signal baselines, and to assess the effects of embedding strategy, truncation depth, model choice, and sensor configuration. Experiments were conducted on the UCI HAR dataset using a fully reproducible pipeline with the original train--test split preserved and subject-disjoint validation to prevent leakage. Three representations were compared: raw flattened signals, time-augmented paths, and lead--lag transformed paths. Signature features were computed at multiple truncation depths and evaluated using multilayer perceptron (MLP) and Random Forest (RF) classifiers under identical preprocessing and validation procedures. A prior K-means-based feature reduction study was also reproduced for comparison. Signature-based representations improved performance when paired with RF models, the best configuration was time-augmented six-channel signatures at depth 6 using entropy-based RF achieving 0.858 accuracy and 0.859 macro F1, outperforming the strongest raw baseline (0.816 accuracy). Lead--lag representations were competitive at moderate depths but did not surpass the best time-augmented models. MLP models did not exceed raw baselines. Signature-based feature learning can improve HAR, but its benefit depends on alignment between representation design and classifier choice.
发表机构
- CNRS(法国国家科学研究中心)
机构由 AI 辅助整理,请以论文原文为准。