发表机构
Johns Hopkins University(约翰斯·霍普金斯大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对视觉HOI识别的隐私与光照问题,提出仅用射频信号的RF-HOI框架,融合毫米波雷达与RFID,结合合成RF数据提升性能,效果优于基线且接近视觉模型。
AI 中文摘要
人体-物体交互(HOI)识别对智能系统至关重要,是虚拟现实、增强现实、具身智能和辅助机器人等应用的基础。然而,基于视觉的HOI方法面临隐私问题和光照条件差的挑战。本研究提出RF-HOI,这是首个仅利用射频(RF)信号进行HOI识别的框架。RF-HOI的一个关键挑战是单模态RF传感无法同时识别动作和交互对象,RF-HOI通过一种新颖的模态融合方法解决该问题,融合毫米波雷达与RFID,实现动作识别与目标识别同步进行。另一个挑战是不同场景下训练数据有限,会降低识别模型的泛化能力,为克服这一点,我们开发了一个模拟器,可大规模合成多样化HOI的多模态RF数据,使我们仅用少量真实数据即可进行微调。实验结果表明,RF-HOI的性能优于所有基线方法,接近视觉模型的性能,且我们的多样化合成训练数据能显著提升系统在真实场景中的性能。这些结果凸显了多模态RF传感在实现鲁棒且隐私保护的HOI识别方面的潜力,以及我们的RF数据合成方法的有效性。
英文摘要
Recognizing Human-Object Interactions (HOI) is essential for intelligent systems, underpinning applications in virtual and augmented reality, embodied AI, and assistive robotics. However, vision-based HOI methods face challenges in privacy concerns and poor light conditions. In this work, we introduce RF-HOI, the first framework that only uses radio frequency (RF) signals for HOI recognition. A key challenge of RF-HOI is that single-modality RF sensing is insufficient to recognize both actions and the objects being interacted with. RF-HOI addresses this through a novel modality fusion that combines mmWave radar and RFID, enabling simultaneous action recognition and target identification. Another challenge is limited training data across diverse setups, which impairs the generalizability of the recognition model. To overcome this, we develop a simulator that synthesizes multimodal RF data for diverse HOIs at scale, allowing us to fine-tune with only a small amount of real-world data. Experiment results show that RF-HOI outperforms all baselines, approaching vision model performance, and that our diverse synthetic training data can significantly boost our system's performance on real-world scenarios. These results highlight the potential of multimodal RF sensing for robust and privacy-preserving HOI recognition as well as the effectiveness of our RF data synthesis.
CommentsAccepted by ACM IMWUT