arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

WiFi-JEPA:用于WiFi-CSI 3D人体姿态估计的自监督学习

WiFi-JEPA: Self-supervised Learning for WiFi-CSI 3D Human Pose Estimation

Doeon Kim, Jungyoon Lee, Seongsin Kim, Seong-heum Kim

arXiv 2607.11064首次发表:更新:

发表机构

Soongsil University(崇实大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对现有WiFi姿态估计器的问题,提出自监督框架WiFi-JEPA。通过特定分词和掩码、模拟管道生成数据,实现CSI原生表示学习。在Person-in-WiFi-3D上成果显著,优于基线,还证明模拟CSI作用及其他目标的不足,提升了下游姿态估计性能。

AI 中文摘要

WiFi信道状态信息(CSI)能够在无摄像头环境中实现隐私保护的人体姿态感知,但现有的基于WiFi的姿态估计器在环境变化时往往失效,且依赖成本高昂的基于摄像头的标注流程,限制了规模。我们提出了WiFi-JEPA,这是一个自监督框架,通过预测掩码潜在嵌入来学习CSI原生表示,而非重建可能包含硬件特定伪像的原始CSI信号。WiFi-JEPA有三点贡献:一是针对CSI张量在信道、时间和链路(C、T、L)上进行特定的CSI分词和链路掩码;二是一个光线追踪CSI模拟管道,可从随机几何原语生成多样的未标注CSI,提供无姿态标注的可扩展预训练数据;三是在Person-in-WiFi-3D上取得了领先成果,在相同评估协议下,WiFi-JEPA在单人和多人3D姿态估计方面均优于先前的WiFi-CSI基线。我们还表明,模拟CSI为真实CSI提供了补充预训练信号,并且四种视觉原生自监督学习目标会使性能下降,而WiFi-JEPA持续提升下游姿态估计性能。

英文摘要

WiFi Channel State Information (CSI) enables privacy-preserving human pose sensing in camera-denied environments, but existing WiFi-based pose estimators often fail under environment shifts and rely on costly camera-based annotation pipelines that limit scale. We propose WiFi-JEPA, a self-supervised framework that learns CSI-native representations by predicting masked latent embeddings instead of reconstructing raw CSI signals that may contain hardware-specific artifacts. WiFi-JEPA makes three contributions: (i) CSI-specific tokenization and link masking tailored to the CSI tensor over channel, time, and link (C,T,L); masking entire Tx-Rx antenna links forces the model to predict one spatial link view from others, capturing cross-link correlations informative of 3D spatial structure. (ii) A ray-tracing CSI simulation pipeline that generates diverse unlabeled CSI from randomized geometric primitives, providing scalable pre-training data without pose annotations. (iii) State-of-the-art results on Person-in-WiFi-3D: WiFi-JEPA outperforms prior WiFi-CSI baselines on both single- and multi-person 3D pose estimation under the same evaluation protocol. We also show that simulated CSI provides complementary pre-training signal to real CSI, and that four vision-native SSL objectives degrade performance below training from scratch, whereas WiFi-JEPA consistently improves downstream pose estimation.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑