发表机构
Singapore Institute of Technology; Singapore University of Technology and Design; University of Toronto(新加坡理工大学; 新加坡科技设计大学; 多伦多大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
WiSPER通过姿态感知预测预训练和条件残差流细化,利用WiFi CSI实现多人3D姿态估计,在PiW3D数据集上显著降低误差。
AI 中文摘要
利用WiFi信道状态信息(CSI)进行多人3D姿态估计具有挑战性,因为不同人的反射信号重叠,无法直接识别各个关节。现有的掩码嵌入目标在没有显式姿态监督的情况下捕获无线关系,而结构化解码器可能保留坐标误差。我们提出了WiSPER,一个两阶段框架,结合了姿态感知的预测预训练与条件残差流细化。姿态感知掩码嵌入学习(PAMEL)将掩码潜在预测与同一CSI上下文上的辅助姿态集监督相结合,引导编码器从部分观测中进行关节定位。基于Transformer的残差流细化(ReFT)生成一组姿态候选以容纳可变人数,并通过由粗略坐标和逐关节解码器特征引导的条件流细化每个候选。两个阶段在训练期间使用配对的CSI和姿态标注,而推理仅需要CSI。在PiW3D数据集上的实验表明,WiSPER实现了总体平均每关节位置误差63.72毫米,相对于WiFi-JEPA降低了40.0%。对于两人和三人实验,WiSPER分别将MPJPE降低了42.1%和38.1%。姿态监督的预训练配置比仅CSI的JEPA获得更低的误差,启用训练好的残差细化器在评估配置中将总体MPJPE降低了13.8-15.6%。
英文摘要
Multi-person 3D pose estimation with WiFi channel state information (CSI) is challenging because reflections from different people overlap without directly identifying individual joints. Existing masked embedding objectives capture wireless relationships without explicit pose supervision, while structured decoders can retain coordinate errors. We propose WiSPER, a two-stage framework combining pose-aware predictive pretraining with conditional residual flow refinement. Pose-Aware Masked Embedding Learning (PAMEL) couples masked latent prediction with auxiliary pose-set supervision on the same CSI context, guiding the encoder toward joint localization from partial observations. Residual Flow refinement with Transformer (ReFT) generates a set of pose candidates to accommodate a variable number of people and refines each candidate through a conditional flow guided by its coarse coordinates and per-joint decoder features. Both stages use paired CSI and pose annotations during training, while inference requires only CSI. Experiments on the PiW3D dataset show that WiSPER achieves an overall mean per-joint position error of 63.72 mm, a 40.0% reduction relative to WiFi-JEPA. For experiments with two and three people, WiSPER reduces MPJPE by 42.1% and 38.1%, respectively. Pose-supervised pretraining configurations obtain lower errors than CSI-only JEPA, and enabling the trained residual refiner reduces overall MPJPE by 13.8-15.6% across the evaluated configurations.