arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

在重要时刻思考:用于社交导航的基于RL策略的条件VLM推理

Think When It Matters: Conditional VLM Reasoning for Social Navigation with RL Policies

Ali Ahmadi, Hamed Rahimi, Adrien Jacquet Cretides, Marie Samson, Mahdi Khoramshahi, Mohamed Chetouani

arXiv 2607.10991首次发表:更新:

发表机构

Institut des Systèmes Intelligents et de Robotique (ISIR)(智能系统与机器人研究所)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究社交机器人导航中强化学习策略缺乏语义推理能力的问题,提出HUMA混合架构,动态平衡RL策略与VLM的优势。在基准测试中任务成功率提高,减少碰撞,消融研究验证组件,实际部署证明方法可行。

AI 中文摘要

随着移动机器人更多地融入日常人类环境,社交机器人导航对于确保人类舒适度、安全性和信任至关重要。强化学习(RL)导航策略虽能提供实时部署所需的快速推理和反应行为,但缺乏灵活语义推理能力,难以推广到复杂社交场景。近期方法转向视觉语言模型(VLM)以改善机器人导航中的语义和社交推理,但其高计算成本和慢推理阻碍实时部署。为克服这些限制,我们引入HUMA(多模态社交导航混合理解),一种动态平衡RL策略计算效率与VLM深度语义理解的混合架构。我们的方法使用反应式RL策略处理低密度、常规导航任务,当人类进入敏感情况(如机器人接近区域)时,以经过后训练的高级VLM为条件。我们在Social-MP3D和Social-HM3D基准上评估HUMA,分别实现任务成功率提高20%和3%,同时显著减少个人空间侵犯和与人类的碰撞。广泛的消融研究验证了每个架构组件,在Mirokaï移动机器人上的实际部署进一步证明了我们方法的实际可行性。

英文摘要

As mobile robots become more integrated into everyday human environments, social robot navigation is becoming essential for ensuring human comfort, safety, and trust. While reinforcement learning (RL) navigation policies provide the fast inference and reactive behavior necessary for real-time deployment, they still lack flexible semantic reasoning capabilities and often fail to generalize to complex social scenarios. Recent approaches have increasingly turned to vision-language models (VLMs) in place of RL policies to improve semantic and social reasoning in robot navigation. Nevertheless, their high computational cost and slow inference remain major barriers to real-time deployment. To overcome these limitations, we introduce HUMA (Hybrid Understanding for Multi-modal social Navigation), a hybrid architecture that dynamically balances the computational efficiency of RL policies with the deep semantic understanding of VLMs. Our approach uses a reactive RL policy to handle low-density, routine navigation tasks, while conditioning it on a post-trained high-level VLM when a human enters sensitive situations, such as the robot's proximity zone. We evaluate HUMA on the Social-MP3D and Social-HM3D benchmarks, where it achieves task success improvements of 20% and 3%, respectively, while significantly reducing personal space violations and human collisions against state-of-the-art baselines. Extensive ablation studies validate each architectural component, and real-world deployment on the Mirokaï mobile robot further demonstrates the practical viability of our approach.

CommentsCoRL 2026 submission. 15 pages, 7 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑