学习多楼层物体导航的模块化策略:一个用于诊断研究的因子化框架
Learning Modular Policy for Multi-Floor Object Navigation:A Factorized Framework for Diagnostic Study
浏览论文内容
中文总结 AI 辅助
针对多楼层物体导航中稀疏奖励的挑战,提出因子化模块化策略,分解为楼层内探索与楼层间切换,并利用VLM蒸馏初始化,实验揭示感知和楼梯攀爬是主要瓶颈。
中文摘要 AI 辅助
多楼层场景中的物体目标导航(ObjectNav)由于长时程决策导致的稀疏奖励而面临挑战。本文提出了一项基于模块化框架的诊断研究,该框架采用有效的可学习策略来分析多楼层场景中的失败因素。为了获得有效的诊断策略,我们设计了层次化因子分解策略,将单一全局策略分解为楼层内探索策略和楼层间切换策略。为了为强化学习(RL)提供有效的初始化,轻量级楼层内策略通过蒸馏视觉语言模型(VLMs)的探索逻辑来学习。在理想化假设下,我们证明了因子化策略在策略表示层面理论上等价于单一全局策略。实验结果表明,感知性能和楼梯攀爬稳定性是多楼层导航中的主要瓶颈。
英文摘要
Object-goal navigation (ObjectNav) in multi-floor scenarios presents a challenge due to sparse rewards caused by long-horizon decision-making. In this paper, we propose a diagnostic study based on a modular framework with an effective learnable policy to analyze failure factors in multi-floor scenarios. To achieve an effective policy for diagnosis, we design the hierarchical factorization policy that deconstructs a single global policy into an intra-floor exploration policy and an inter-floor switching policy. To providing an effective initialization for Reinforcement Learning (RL), the lightweight intra-floor policy is learned by distilling the exploration logic of Visual Language Models (VLMs). Under idealized assumptions, we show that the factorized policy is theoretically equivalent to a single global policy at the policy-representation level. Experiment results indicate that perception performance and stair climbing stability are the primary bottlenecks in multi-floor navigation.
发表机构
- State Key Laboratory of Industrial Control and Technology, Zhejiang University(浙江大学工业控制技术国家重点实验室)
机构由 AI 辅助整理,请以论文原文为准。