arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.06958cs.RO

学习多楼层物体导航的模块化策略:一个用于诊断研究的因子化框架

Learning Modular Policy for Multi-Floor Object Navigation:A Factorized Framework for Diagnostic Study

Shichao Zhai, Shuhao Ye, Rong Xiong, Yue Wang

首次发表
浏览论文内容

中文总结 AI 辅助

针对多楼层物体导航中稀疏奖励的挑战,提出因子化模块化策略,分解为楼层内探索与楼层间切换,并利用VLM蒸馏初始化,实验揭示感知和楼梯攀爬是主要瓶颈。

中文摘要 AI 辅助

多楼层场景中的物体目标导航(ObjectNav)由于长时程决策导致的稀疏奖励而面临挑战。本文提出了一项基于模块化框架的诊断研究,该框架采用有效的可学习策略来分析多楼层场景中的失败因素。为了获得有效的诊断策略,我们设计了层次化因子分解策略,将单一全局策略分解为楼层内探索策略和楼层间切换策略。为了为强化学习(RL)提供有效的初始化,轻量级楼层内策略通过蒸馏视觉语言模型(VLMs)的探索逻辑来学习。在理想化假设下,我们证明了因子化策略在策略表示层面理论上等价于单一全局策略。实验结果表明,感知性能和楼梯攀爬稳定性是多楼层导航中的主要瓶颈。

英文摘要

Object-goal navigation (ObjectNav) in multi-floor scenarios presents a challenge due to sparse rewards caused by long-horizon decision-making. In this paper, we propose a diagnostic study based on a modular framework with an effective learnable policy to analyze failure factors in multi-floor scenarios. To achieve an effective policy for diagnosis, we design the hierarchical factorization policy that deconstructs a single global policy into an intra-floor exploration policy and an inter-floor switching policy. To providing an effective initialization for Reinforcement Learning (RL), the lightweight intra-floor policy is learned by distilling the exploration logic of Visual Language Models (VLMs). Under idealized assumptions, we show that the factorized policy is theoretically equivalent to a single global policy at the policy-representation level. Experiment results indicate that perception performance and stair climbing stability are the primary bottlenecks in multi-floor navigation.

发表机构

  • State Key Laboratory of Industrial Control and Technology, Zhejiang University(浙江大学工业控制技术国家重点实验室)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑