arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.21008cs.RO

SPARROW:用于自适应机器人路径规划、观测与等待的Survival-POMCP方法

SPARROW: Survival-POMCP for Adaptive Robot Routing, Observation, and Waiting

Hshmat Sahak, Aoran Jiao, Nicholas Rhinehart, Timothy D. Barfoot

首次发表
浏览论文内容

中文总结 AI 辅助

SPARROW提出基于POMCP的信念空间规划器,结合生存模型与学习价值准则,处理临时障碍下的导航决策,相比OSCAR降低平均到达时间12%-26%。

中文摘要 AI 辅助

可能阻塞机器人预定路线的临时障碍物会引发一个序列导航问题:机器人必须决定是等待障碍物清除、重新规划路线,还是在行动前获取更多关于障碍物的信息。我们将临时障碍物间的图导航问题建模为部分可观测半马尔可夫决策过程,并引入SPARROW——一种基于部分可观测蒙特卡洛规划(POMCP)构建的信念空间规划器。SPARROW在遍历、观测和有限时长等待动作之间进行搜索,同时维护关于潜在障碍物类别和清除时间的粒子信念。类别条件生存模型从清除观测和右删失遭遇(即机器人在观察到清除前重新规划路线的情况)中在线学习。生成模型在每次动作展开时模拟障碍物的到达和清除,使规划器能够考虑替代路线上可能出现的阻塞。我们进一步引入一个学习价值准则,权衡收集带标签生存数据的即时成本与其对未来导航遗憾的预期减少。在两个仿真图和多障碍物类别设置中,SPARROW相对于OSCAR(一种针对同一问题的近期基于生存的方法)将平均到达目标时间降低了12%-26%。在物理移动机器人上,SPARROW相对于OSCAR将平均到达目标时间降低了20.5%,同时根据环境条件变化选择性地进行观测、等待和重新规划路线。

英文摘要

Temporary obstacles that may block a robot's planned route create a sequential navigation problem: a robot must decide whether to wait for a blockage to clear, reroute, or acquire more information about the obstacle before acting. We formulate graph navigation among temporary obstacles as a partially observable semi-Markov decision process and introduce SPARROW, a belief-space planner built on Partially Observable Monte Carlo Planning (POMCP). SPARROW searches over traversal, observation, and finite-duration waiting actions while maintaining a particle belief over latent obstacle classes and clearance times. Class-conditioned survival models are learned online from both clearance observations and right-censored encounters where the robot reroutes before clearance is observed. A generative model simulates obstacle arrivals and clearances as each action unfolds, so the planner can account for blockages that may occur along alternative routes. We further introduce a value-of-learning criterion that trades the immediate cost of collecting labelled survival data against its expected reduction in future navigation regret. Across two simulation graphs and multiple obstacle-class settings, SPARROW reduces mean time-to-goal by 12-26% relative to OSCAR, a recent survival-based method for the same problem. On a physical mobile robot, SPARROW reduces mean time-to-goal by 20.5% relative to OSCAR while selectively observing, waiting, and rerouting as environment conditions change.

发表机构

  • University of Toronto Institute of Aerospace Studies (UTIAS)(多伦多大学航空航天研究所)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑