arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

面向广义任务与运动规划问题的编码智能体

Coding Agents for Generalized Task and Motion Planning Problems

Matteo Merler, Bowen Li, Josh Roy, Yichao Liang, Qianwei Wang, Yixuan Huang, Tom Silver

arXiv 2609.30233首次发表:更新:

发表机构

Fondazione Bruno Kessler; Carnegie Mellon University; Princeton University; University of Cambridge(布鲁诺·凯斯勒基金会; 卡内基梅隆大学; 普林斯顿大学; 剑桥大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文研究编码智能体能否通过合成程序自动化广义任务与运动规划,实验表明其成功率优于手工规划器,且计算量低一个数量级。

AI 中文摘要

即使在完全可观测且具有对象中心状态的情况下,任务与运动规划(TAMP)问题仍然困难,因为离散决策与几何、运动学和动力学约束紧密耦合。广义TAMP通过利用跨问题实例的规律性来减少新实例上的规划工作量,从而解决这一困难。然而,现有方法需要大量的TAMP特定工程。我们研究编码智能体是否可以通过合成跨实例泛化的程序来自动化这一过程。给定任务描述和模拟器访问权限,每个智能体在固定的合成预算内开发程序的同时,选择如何与环境交互。程序随后被冻结,并在未见过的实例上进行评估。我们在来自KinDER和PDDLStream的28个模拟环境中评估了Claude Code (Opus 5)和Codex (GPT-5.6 Sol和GPT-6 Astra),这些环境中的对象数量超出了原始基准中评估的数量。在所有程序合成方法中,我们在100个保留实例上评估了980个生成的程序,每个实例进行100次评估,总计98,000次评估回合。总体而言,我们发现编码智能体在广义TAMP上出奇地有效:在16个有规划器可用的环境中,所有三种智能体配置在平均成功率上均优于手工设计的规划器、一次性生成和基于LLM的广义规划基线(智能体为56%至95%,而规划器为47%)。随着对象数量的增长,智能体的程序比规划器保持更高的成功率,每个实例平均使用的计算量少一个数量级。日志显示智能体使用交互来校准物理模型、测试边缘情况并完善策略。我们发布了所有代码,包括提供给智能体的完整提示。这些发现表明,编码智能体是广义TAMP的一个强基线。

英文摘要

Task and motion planning (TAMP) problems remain difficult even with full observability and object-centric states because discrete decisions are tightly coupled to geometric, kinematic, and dynamic constraints. Generalized TAMP addresses this difficulty by exploiting regularities across problem instances to reduce planning effort on new instances. However, existing methods require substantial TAMP-specific engineering. We investigate whether coding agents can automate this process by synthesizing programs that generalize across instances. Given a task description and simulator access, each agent chooses how to interact with the environment while developing a program within a fixed synthesis budget. The program is then frozen and evaluated on unseen instances. We evaluate Claude Code (Opus 5) and Codex (GPT-5.6 Sol and GPT-6 Astra) on 28 simulated environments from KinDER and PDDLStream, with object counts beyond those evaluated in the original benchmark. Across all program synthesis methods, we evaluate 980 generated programs on 100 held-out instances each, 98,000 evaluation episodes in total. Overall, we find that coding agents are surprisingly effective at generalized TAMP: all three agent configurations outperform hand-engineered planners, one-shot generation, and an LLM-based generalized planning baseline in mean success (56% to 95% versus 47% for the planners, on the 16 environments where a planner is available). As object counts grow, the agents' programs maintain higher success than the planner, using an order of magnitude less computation per instance on average. Logs show agents using interaction to calibrate physical models, test edge cases, and refine strategies. We release all code, including the full prompts given to the agents. These findings suggest that coding agents are a strong baseline for generalized TAMP.

Comments9 pages, 4 figures, 3 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑