arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于网格基础模型监督的毫米波雷达隐私保护全身网格重建

Privacy-Preserving Full-Body Meshing from mmWave Radar via Mesh Foundation Model Supervision

Shuxing Zhang, Yongquan Ni, Zhenyu Ding, Yawen Lin

arXiv 2609.34768首次发表:更新:

发表机构

AI Value Center, Incaier (Haier)(AI价值中心,因凯尔(海尔))

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出跨模态师生框架,利用SAM 3D Body教师生成全身网格真值,结合StudentPoseFormer实现商用毫米波雷达的全身3D网格重建,在MM-Fi基准达7.45 cm MPJPE,推理仅需雷达。

AI 中文摘要

毫米波(mmWave)雷达能够实现隐私保护的人体感知,但商用单芯片传感器点云的极度稀疏性(平均每帧约6.5个点;约28%的帧为空)将先前的工作局限于身体部位关键点或离散动作分类。我们提出了一种跨模态师生框架,将商用雷达提升至全身、逐帧、公制3D网格重建,并带有逐关节不确定性。三项创新:(1)网格基础模型教师——SAM 3D Body从单个RGB帧生成全身MHR真值(70个关节,18,439个网格顶点),无需训练,将标注成本降低数个数量级;(2)StudentPoseFormer——采用掩码注意力池化的集合编码、时间Transformer和CVAE多假设头,输出姿态均值和逐关节方差,诚实报告雷达无法观测的区域;(3)多阶段真值质量流水线(置信度门控、深度验证、时间平滑、骨骼长度一致性、坏帧剔除)以及系统性的信息利用消融实验。在公共MM-Fi基准(相同TI IWR6843传感器,跨受试者)上,我们的完整配置达到7.45 cm的12关节MPJPE,消融实验证明了点累积(k=3,-0.34 cm)、多普勒(-0.85 cm;快速动作时手腕处-2 cm)和速度损失(-0.27 cm)的因果价值。在我们自有的同步雷达+RGB-D语料库上,采用块级留出划分,流水线实现21.47 cm的端到端误差(逐关节层级从髋部4.8 cm到手腕34.7 cm——与物理信息极限匹配),通过约3万个多样化样本可提升至15 cm,且缩放定律表明样本多样性而非数量是约束因素。部署推理仅依赖雷达——无相机,无图像。

英文摘要

Millimeter-wave (mmWave) radar enables privacy-preserving human perception, but the extreme sparsity of point clouds from commercial single-chip sensors (mean ~6.5 points/frame; ~28% empty frames) has confined prior art to body-part keypoints or discrete action classification. We present a cross-modal teacher-student framework that lifts commercial radar to full-body, per-frame, metric 3D mesh reconstruction with per-joint uncertainty. Three innovations: (1) a mesh-foundation-model teacher - SAM 3D Body produces whole-body MHR ground truth (70 joints, 18,439 mesh vertices) from a single RGB frame with zero training, slashing annotation cost by orders of magnitude; (2) StudentPoseFormer - set encoding with masked attention pooling, a temporal Transformer, and a CVAE multi-hypothesis head that outputs both the pose mean and per-joint variance, honestly reporting where the radar cannot see; and (3) a multi-stage ground-truth quality pipeline (confidence gating, depth validation, temporal smoothing, bone-length consistency, bad-frame rejection) plus systematic information-lever ablations. On the public MM-Fi benchmark (same TI IWR6843 sensor, cross-subject), our full configuration reaches 7.45 cm 12-joint MPJPE, with ablations proving the causal value of point accumulation (k = 3, -0.34 cm), Doppler (-0.85 cm; -2 cm at the wrist on fast actions), and velocity loss (-0.27 cm). On our own synchronized radar + RGB-D corpus with block-level held-out splits, the pipeline achieves 21.47 cm end-to-end (per-joint hierarchy from 4.8 cm at the hip to 34.7 cm at the wrist - matching physical information limits), could be improved to 15 cm with ~30k diverse samples, and a scaling law shows sample diversity, not volume, is the binding constraint. Deployment inference is radar-only - no camera, no image.

CommentsWithdrawn by the authors: the author team is still finalizing the scope and release timing of this work, and will resubmit after internal review

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑