arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

mmHRI:基于毫米波雷达的隐私保护人机交互

mmHRI: Towards Privacy-Preserving Human-Robot Interaction with Millimeter-Wave Radar

Junqiao Fan, Yuxuan Hu, Bofan Lyu, Yanshuo Lu, Pengfei Liu, Jiarui Zhang, Fangqiang Ding, Lihua Xie, Gen Li, Jianfei Yang

arXiv 2609.34220首次发表:更新:

发表机构

Nanyang Technological University; The Hong Kong University of Science and Technology (Guangzhou); Hunan University(南洋理工大学; 香港科技大学(广州); 湖南大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对现有HRI依赖摄像头导致的隐私问题,提出基于毫米波雷达的mmHRI框架,通过双流架构和记忆状态空间模型实现隐私保护的人机交互,在动作识别和遮挡下递送任务中表现优异。

AI 中文摘要

辅助机器人在越来越多的人类中心环境中运行,并执行各种人机交互(HRI)任务,例如物体递送。然而,现有的大多数HRI系统依赖RGB摄像头持续观察人类,以响应非语言指令(如手势)。这在隐私关键环境中引发了隐私担忧,例如医院病房或餐厅,在这些环境中直接对人类的摄像头观察受到限制。为了开发隐私保护的HRI,我们利用毫米波(mmWave)雷达,它能够在不产生可识别图像的情况下感知人类运动,从而穿过隐私屏障。我们提出了mmHRI,这是第一个实现毫米波雷达引导的隐私保护HRI的多模态机器人操作框架。mmHRI引入了两个关键设计,以缓解杂乱机器人操作环境中雷达数据的稀疏性和时间不一致性。首先,我们提出了一种双流架构,联合学习未过滤的原始雷达张量和雷达点云,以估计人类动作和3D姿态。为了缓解信号不一致性,mmHRI进一步整合了基于记忆的状态空间模型(MSSM),该模型保留历史雷达特征以减少姿态/动作的突变。这些估计的人类状态随后被转换为结构化的文本机器人指令,这些指令控制视觉-语言-动作(VLA)策略,用于闭环机器人操作和人类感知反应。我们的评估涵盖人类动作识别和闭环递送与取回。在隐私保护的窗帘设置中,mmHRI实现了85.09%的动作识别准确率,优于现有的基于雷达的替代方案。机器人试验进一步证明了在视觉遮挡下的成功递送和取回,在未见过的受试者、杂乱配置和环境中的任务性能稳定。

英文摘要

Assistive robots increasingly operate in many human-centered environments and perform various human-robot interaction (HRI) tasks, such as object delivery. However, most existing HRI systems rely on RGB cameras that continuously observe humans to respond to non-verbal commands, such as hand gestures. This raises privacy concerns in privacy- critical environments, such as hospital wards or restaurants, where direct camera observation of humans is restricted. To develop privacy-preserving HRI, we leverage millimeter-wave (mmWave) radar, which can sense human motion through privacy barriers without identifiable imagery. We propose mmHRI, the first multi-modal robot manipulation framework that achieves mmWave radar-guided privacy-preserving HRI. mmHRI introduces two key designs to mitigate the sparsity and temporal inconsistency of radar data in cluttered robot manipulation environments. First, we propose a dual-stream architecture that jointly learns from unfiltered raw radar tensors and radar point clouds to estimate both human actions and 3D poses. To mitigate signal inconsistency, mmHRI further incorporates a memory-based state-space model (MSSM) that retains historical radar features to reduce abrupt changes in pose/action. These estimated human states are then converted into structured textual robot instructions, which control a vision-language-action (VLA) policy for closed-loop robot manipulation and human-aware reactions. Our evaluation covers human action recognition and closed-loop delivery and retrieval. In the privacy-preserving curtain setting, mmHRI achieves 85.09% action-recognition accuracy, outperforming existing radar-based alternatives. Robot trials further demonstrate successful delivery and retrieval under visual occlusion, with stable task performance across unseen subjects, clutter configurations, and environments.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑