arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.23629cs.ROcs.AI

TAMP算子学习的宏算子生成与谓词选择

Macro-Operator Generation and Predicate Selection for TAMP Operator Learning

Can Emir Bora, Emre Ugur

首次发表
浏览论文内容

中文总结 AI 辅助

本文针对TAMP算子学习的效率问题,提出自动生成宏算子并剪枝未被引用谓词的系统,在四个TAMP领域实现最高4.6倍规划加速,还解决了基线无法处理的长序列任务。

中文摘要 AI 辅助

手动创建符号算子是部署任务与运动规划(Task and Motion Planning, TAMP)系统的主要瓶颈之一。近期研究表明,这些算子可直接从演示数据中学习。然而现有方法通常孤立学习每个动作,无法捕捉操纵任务中反复出现的多步结构,导致在长序列任务上搜索变得难以处理;符号状态层面还存在另一效率问题:每个提供的谓词会在所有搜索节点处被评估,即便它从未出现在任何学习到的算子中。本文提出一个同时解决这两个问题的系统,其核心组件是自动生成宏算子——将反复出现的单个动作序列压缩为单个规划步骤的复合动作。该系统直接从训练数据中发现因果关联的动作对(即一个动作恰好产生下一动作所需条件的动作对),并将每对转化为新算子;同时,系统会剪枝所有未被任何学习到的算子引用的谓词,从而缩小在每个搜索节点处评估的符号状态。这些改动共同缩短了有效规划 horizon,且带来的益处随任务长度增加而增长。在四个 TAMP 领域中,本文方法与基线方法 Learning Operators for TAMP 相比,实现了最高达 4.6 倍的规划加速;更重要的是,它解决了基线方法无法解决的长序列任务。因此,宏算子发现不仅能加速规划,在某些领域还决定了实际中的可解性。

英文摘要

Creating symbolic operators by hand is one of the main bottlenecks in deploying Task and Motion Planning systems (TAMP). Recent works show that these operators can instead be learned directly from demonstration data. Existing methods, however, typically learn each action in isolation and cannot capture the recurring multi-step structure of manipulation tasks, so the search becomes intractable on long sequential tasks. A further inefficiency arises in the symbolic state: every provided predicate is evaluated at every search node, even when it never appears in any learned operator. We present a system that addresses both problems together. Its central component is the automatic generation of macro-operators, composite actions that compress a recurring sequence of individual actions into a single planning step. Our system discovers causally linked action pairs directly from the training data, where one action produces exactly the condition that the next one requires, and turns each pair into a new operator. Alongside this, our system prunes every predicate that no learned operator references, which shrinks the symbolic state evaluated at each search node. Together, these changes shorten the effective planning horizon, and the benefit they bring grows with the length of the task. Across four TAMP domains, our method reaches up to a 4.6x planning speedup compared to the baseline method, namely Learning Operators for TAMP. More importantly, it solves a long sequential task that the baseline cannot solve. Macro-operator discovery thus not only accelerates planning but, in certain domains, determines solvability in practice.

发表机构

  • Bogazici University(博加齐奇大学)

机构由 AI 辅助整理,请以论文原文为准。

↑