arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.04127cs.CV

教基础模型阅读毫米波:用于人类行为理解的姿态引导运动学表示

Teaching Foundation Models to Read mmWave: Pose-Guided Kinematic Representation for Human Behavior Understanding

Duo Zhang, Zhehui Yin, Zhiyun Yao, Haotong Qin, Xusheng Zhang, Hongliu Yang, Jianyu Sun, Junzhe Wang, Zizhou Fan, Michele Magno, Daqing Zhang

首次发表
浏览论文内容

中文总结 AI 辅助

本研究提出雷达-语言模型mmMind,利用同步3D姿态监督训练,结合时空雷达编码器与LLM,构建含17.9小时数据的mmMind-Bench,在多项任务上优于现有基线。

中文摘要 AI 辅助

大型语言模型智能体需要感知物理环境中的人类行为,毫米波(mmWave)雷达提供了一种隐私友好且非接触的感知方式,但雷达观测结果难以与语言对齐。现有的雷达-语言方法常依赖合成数据,或缺乏对人体结构与运动的显式监督。我们提出mmMind,一种雷达-语言模型,使用同步的3D姿态作为仅用于训练的监督信号。我们预训练了一个时空雷达编码器以捕获身体构型与运动动力学,之后移除姿态头,使得推理仅需雷达数据。随后将学习到的雷达表示与大型语言模型(LLM)对齐,用于行为描述和时空问答。我们还引入了mmMind-Bench,一个真实世界的雷达-语言基准,包含23名参与者在7个室内环境中录制的17.9小时数据。在描述、问答及未见动作泛化任务上的实验显示,mmMind始终优于现有的雷达-语言基线,而 ablation( ablation 研究)证实了姿态引导预训练的重要性。

英文摘要

Large language model agents need to perceive human behavior in physical environments. Millimeter-wave (mmWave) radar provides a privacy-friendly and contactless sensing modality, but radar observations are difficult to align with language. Existing radar-language methods often rely on synthetic data or lack explicit supervision for human body structure and motion. We present mmMind, a radar-language model that uses synchronized 3D pose as training-only supervision. A spatio-temporal radar encoder is pretrained to capture body configuration and motion dynamics, after which the pose head is removed so that inference requires radar alone. The learned radar representations are then aligned with an LLM for behavior captioning and spatio-temporal question answering. We also introduce mmMind-Bench, a real-world mmWave-language benchmark containing 17.9 hours of recordings from 23 participants across seven indoor environments. Experiments on captioning, question answering, and unseen-action generalization show that mmMind consistently outperforms existing radar-language baselines, while ablations confirm the importance of pose-guided pretraining.

补充信息

↑