arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

全模态分解自编码器学习全栈可穿戴解耦表示

Omni-modal decomposition autoencoders learn full-stack wearable disentangled representations

Ioannis Ziogas, Ensieh Khazaei, Bilal Taha, Aamna Al Shehhi, Ahsan H. Khandoker, Leontios J. Hadjileontiadis, Dimitrios Hatzinakos

arXiv 2608.07385首次发表:更新:

发表机构

Khalifa University; University of Toronto; MIT Media Lab; Aristotle University of Thessaloniki(哈利法大学; 多伦多大学; 麻省理工学院媒体实验室; 亚里士多德大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对现有多模态可穿戴模型的不足,提出OmniDecVAEs框架,在30模态的HAR任务中,提升了识别准确率与数据合成质量,可用于边缘可穿戴与医疗领域。

AI 中文摘要

学习解耦表示是开发多模态可穿戴计算中通用、多功能且可持续模型的关键要求。然而,现有方法无法作为全栈可穿戴处理器运行,即它们无法同时解决特定任务的分类性能、解耦且可解释的表示学习、融合以及高度异构多模态时间序列的生成建模问题。为解决这一差距,我们引入全模态变分分解自编码器(OmniDecVAEs),这是一个可从任意数量的模态中以统一且可扩展的方式高效学习多用途表示的框架。OmniDecVAEs 通过多视图自监督分解损失和共享非对称自编码器(AE)架构学习模态条件时频潜在子空间,从而扩展了 DecVAEs。在具有多达三十种模态的挑战性全模态人类活动识别(HAR)设置上的结果表明,OmniDecVAEs 具备学习全栈可穿戴表示的能力。与基于 Transformer 和 VAE 的方法相比,OmniDecVAEs 的全栈解耦表示属性在活动和身份识别中分别带来 1.01% 和 6.75% 的准确率提升。此外,OmniDecVAEs 合成逼真的全模态时频数据,表现出增强的重构效果(平均绝对误差提升 76.84%)以及真实数据与合成数据间的分布相似性(最大均值差异提升 13.85%)。我们的结果凸显了 OmniDecVAEs 作为适用于智能边缘可穿戴设备和临床医疗的轻量型模型的潜力,它通过增强的表示能力、模态不变空间复杂度(410 万参数)和实时延迟,在单个模型中统一了处理需求和能力。

英文摘要

Learning disentangled representations is a key requirement for developing versatile, general-purpose, and sustainable models in multi-modal wearable computing. However, existing approaches do not operate as full-stack wearable processors, i.e., they do not simultaneously address task-specific classification performance, disentangled and interpretable representation learning, fusion, and generative modeling of highly heterogeneous multi-modal time series. To address this gap, we introduce Omni-modal Variational Decomposition Autoencoders (OmniDecVAEs), a framework that efficiently learns multi-purpose representations in a unified and scalable manner from arbitrarily many modalities. OmniDecVAEs extend DecVAEs by learning modality-conditioned time-frequency latent subspaces through a multi-view self-supervised decomposition loss and a shared asymmetric autoencoder (AE) architecture. Results on a challenging omni-modal human activity recognition (HAR) setting with up to thirty modalities, demonstrate the ability of OmniDecVAEs to learn full-stack wearable representations. When compared to transformer-based and VAE-based methods, OmniDecVAEs full-stack disentangled representation properties lead to accuracy improvements of 1.01% and 6.75% in activity and identity recognition, respectively. Furthermore, OmniDecVAEs synthesize realistic omni-modal time-frequency data that manifest with enhanced reconstructions (mean absolute error improves by 76.84%) and distributional similarity between real and synthetic data (maximum mean discrepancy improves by 13.85%). Our results highlight OmniDecVAEs potential as a lightweight model suitable for intelligent edge wearables and clinical healthcare, unifying processing requirements and abilities in a single model, through its enhanced representational capacity, modality-invariant spatial complexity (4.1M parameters), and real-time latency.

Comments15 pages, 7 figures, 7 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑