arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Imagine-TAMP:部分可观测性下的想象引导任务与运动规划

Imagine-TAMP: Imagination-Guided Task and Motion Planning in Partial Observability

Antareep Singha, Shivaram Kumar, Yoonwoo Kim, Yoonchang Sung

arXiv 2609.20396首次发表:更新:

发表机构

Nanyang Technological University; The University of Texas at Austin(南洋理工大学; 德克萨斯大学奥斯汀分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

Imagine-TAMP利用语义和几何想象在部分可观测性下评估任务策略,改善观测与操作决策,提升成功率并减少规划时间。

AI 中文摘要

在杂乱环境中运行的机器人通常必须操作位置仅部分可观测的物体。一个核心挑战是决定是获取另一次观测,还是先操作可能遮挡目标的物体。传统的任务与运动规划(TAMP)方法通常使用符号动作成本或昂贵的几何规划来做出这一决策,但这两者都无法充分捕捉观测揭示被遮挡目标的可能性。我们提出了Imagine-TAMP,一个交错式规划与执行框架,在部分可观测性下,利用语义和几何想象来比较替代的任务级策略,然后再投入昂贵的运动规划。视觉-语言模型利用目标与可见物体之间的常识关系,形成目标位置的粒子信念,而生成式场景模型则估计未观测区域中合理的几何结构。给定目标假设和想象场景,Imagine-TAMP生成多个符号规划骨架,并分配非单元成本,这些成本近似于操作努力和来自感知动作的目标可见性,从而区分一个短但信息量不足的观测策略与一个先操作遮挡物以更好暴露目标的较长策略。选定的骨架随后被细化为可行的连续规划并执行,新的观测更新信念,并在必要时触发重新规划。实验表明,想象引导的评估改善了观测与操作之间的决策:在视角受限的货架场景中,非单元几何评估将成功率从46.0%提高到84.0%,而语义信念塑造进一步减少了操作和重新规划。在真实机器人上,完整系统相对于仅几何的消融,将规划时间减少了32%。

英文摘要

Robots operating in cluttered environments must often manipulate objects whose locations are only partially observable. A central challenge is deciding whether to acquire another observation or to first manipulate objects that may occlude the target. Conventional task and motion planning (TAMP) approaches typically make this decision using symbolic action costs or expensive geometric planning, neither of which adequately captures how likely an observation is to reveal an occluded target. We introduce Imagine-TAMP, an interleaved planning and execution framework that uses semantic and geometric imagination to compare alternative task-level strategies under partial observability before committing to expensive motion planning. A vision-language model shapes a particle belief over target locations using commonsense relationships between the target and visible objects, while a generative scene model estimates plausible geometry in unobserved regions. Given a target hypothesis and imagined scene, Imagine-TAMP generates multiple symbolic plan skeletons and assigns non-unit costs that approximate both manipulation effort and target visibility from sensing actions, distinguishing a short but poorly informative observation strategy from a longer strategy that first manipulates an occluder to better expose the target. The selected skeleton is then refined into a feasible continuous plan and executed, with new observations updating the belief and triggering replanning when necessary. Experiments show that imagination-guided evaluation improves observation-versus-manipulation decisions: in viewpoint-constrained shelf scenes, non-unit geometric evaluation increases success from 46.0% to 84.0%, while semantic belief shaping further reduces manipulation and replanning. On a real robot, the complete system reduces planning time by 32% relative to a geometry-only ablation.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑