arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.01351cs.ROcs.AI

面向高维部分可观测马尔可夫决策过程(POMDP)的可扩展Rao-Blackwell化在线规划

Scalable Rao-Blackwellized Online Planning for High-Dimensional POMDPs

  • University of Colorado Boulder(科罗拉多大学博尔德分校)
  • University of Massachusetts Amherst(马萨诸塞大学阿默斯特分校)

机构由 AI 辅助整理,请以论文原文为准。

Jiho Lee, Nisar Ahmed, Kyle Hollins Wray, Zachary Sunberg

中文总结 AI 辅助

本研究扩展RB-POMDP框架,采用混合连续-离散信念表示,结合FastSLAM 2.0,在机器人搜救任务中实现高维POMDP下的高效在线规划,性能优于纯采样方法。

中文摘要 AI 辅助

在部分可观测环境中运行的机器人系统,其高维状态空间下的不确定性在线规划仍是核心挑战。基于采样的POMDP求解器虽能在大型或连续域中实现近似决策,但因蒙特卡洛估计固有的高方差,其性能会随信念维度增加而下降。本研究扩展了Rao-Blackwell化在线POMDP(RB-POMDP)框架,通过混合连续-离散信念表示提升其在高维场景中的泛化性;在基于树的规划过程中,通过解析传播边缘化状态分量相关的不确定性,所提方法降低了值估计中采样导致的方差。我们将该框架与FastSLAM 2.0集成,在机器人搜救任务中验证了其有效性。实验结果表明,在相同计算预算下,所提规划器相比纯采样方法,使用的粒子数和规划模拟次数显著更少,却能获得更高的累积奖励。这些结果表明,在RB-POMDP框架中,可有效利用具有可处理充分统计量的结构化高维机器人问题,实现计算可行的在线决策。

英文摘要

Online planning under uncertainty remains a fundamental challenge for robotic systems operating in partially observable environments with high-dimensional state spaces. While sampling-based POMDP solvers enable approximate decision-making in large or continuous domains, their performance degrades as belief dimensionality increases due to the high variance inherent in Monte Carlo-based estimation. In this work, we extend the Rao-Blackwellized online POMDP (RB-POMDP) framework to improve its generalizability in high-dimensional settings through hybrid continuous-discrete belief representations. By analytically propagating uncertainty associated with marginalized state components during tree-based planning, the proposed approach reduces sampling-induced variance in value estimation. We demonstrate the effectiveness of this framework in a robotic search-and-rescue task by integrating it with FastSLAM 2.0. Experimental results show that the proposed planner achieves higher cumulative rewards using significantly fewer particles and planning simulations than purely sampling-based methods under equivalent computational budgets. These results suggest that structured high-dimensional robotic problems admitting tractable sufficient statistics can be effectively leveraged within the RB-POMDP framework for computationally feasible online decision-making.

↑