SOPD-SocialNav:用于视觉语言社交导航的选择性在线策略蒸馏
SOPD-SocialNav: Selective On-Policy Distillation for Vision-Language Social Navigation
- Graduate School of Information Science and Technology, Hokkaido University(北海道大学信息科学与技术研究生院)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对大规模视觉语言模型难部署、轻量级模型社交推理能力不足的问题,提出SOPD-SocialNav方法,通过基于熵的令牌选择机制及温度控制的散度目标,将社交导航知识从大型教师模型转移到轻量级学生模型,实验证明该方法有效。
AI中文摘要:
视觉语言模型在利用对复杂环境和人类行为的丰富语义理解进行社交机器人导航方面显示出强大潜力。然而,大规模视觉语言模型难以部署在资源受限的机器人平台上,而轻量级视觉语言模型往往缺乏足够的社交推理能力。为解决此问题,我们提出了SOPD-SocialNav,一种选择性在线策略蒸馏(SOPD)方法,将社交导航知识从大型教师视觉语言模型转移到轻量级学生视觉语言模型。SOPD引入了基于熵的令牌选择机制,利用教师的不确定性来识别具有社交信息的决策令牌,同时抑制来自对应于琐碎导航状态的低熵令牌的梯度。然后使用温度控制的 Jensen-Shannon 散度目标来对齐所选令牌上的学生和教师分布。在SNEI和MUSON基准上的实验表明,SOPD在动作预测、感知一致性和推理一致性方面始终优于监督微调、离线策略蒸馏和标准在线策略蒸馏基线。在Scout Mini机器人上的实际部署进一步表明,蒸馏后的模型可以在对话和排队场景中生成更符合社交规范的导航行为。这些结果表明SOPD是构建轻量级但具有社交意识的基于视觉语言模型的导航系统的有效策略。
英文摘要:
Vision-language models have shown strong potential for social robot navigation by leveraging rich semantic understanding of complex environments and human behaviors. However, large scale VLMs are difficult to deploy on resource-constrained robotic platforms, while lightweight VLMs often lack sufficient social reasoning capability. To address this problem, we propose SOPD-SocialNav, a selective on-policy distillation (SOPD) method that transfers social navigation knowledge from a large teacher VLM to a lightweight student VLM. SOPD introduces an entropy-based token selection mechanism that uses teacher uncertainty to identify socially informative decision tokens, while suppressing gradients from low-entropy tokens corresponding to trivial navigation states. A temperature-controlled Jensen-Shannon divergence objective is then used to align the student and teacher distributions on the selected tokens. Experiments on the SNEI and MUSON benchmarks demonstrate that SOPD consistently outperforms supervised fine-tuning, off-policy distillation, and standard on-policy distillation baselines in action prediction, perception consistency, and reasoning consistency. Real-world deployment on a Scout Mini robot further shows that the distilled model can generate more socially appropriate navigation behaviors in conversational and queuing scenarios. These results suggest that SOPD is an effective strategy for building lightweight yet socially aware VLM-based navigation systems.