arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

通过编码代理实现基于契约的行为树合成

Contract-Grounded Behavior Tree Synthesis via Coding Agents

Jonathan Salfity, Robert Blake Anderson, Mitch Pryor

arXiv 2607.12220首次发表:更新:

发表机构

The University of Texas at Austin(德克萨斯大学奥斯汀分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究从自然语言合成机器人行为树时基础设定易出现的问题,提出基于契约的合成架构,编码代理查询服务器获取契约后合成行为树,经实验评估两个大语言模型,结果显示该架构能实现高验证率和成功率,且可转移到物理硬件。

AI 中文摘要

从自然语言合成可部署的机器人行为树需要进行基础设定,以确保生成的行为树仅引用机器人实际能执行的技能。现有的基于大语言模型的行为树合成方法往往将基础设定的责任交给提示作者。当作者不了解机器人能执行哪些技能、这些技能如何参数化或机器人运行时软件如何约束有效的行为树结构时,这会使部署变得脆弱。本文提出了一种基于契约的行为树合成架构,其中编码代理在合成行为树进行验证和执行之前,查询机器人端的模型上下文协议(MCP)服务器,以检索由技能库、允许的行为树操作符和可选的行为树组合模板组成的明确契约。在我们的框架中,非专业操作员在不了解机器人实现细节的情况下发出自然语言命令,而机器人运行时验证门在执行前强制执行正确性。我们在PyRoboSim中的110个模拟任务和物理Husarion Panther机器人上的14个任务中评估了两个大语言模型,一个封闭模型(Sonnet 4.6)和一个较小的开源模型(Gemma4:31b)。结果表明,契约基础设定能够实现近乎完美的行为树验证和高任务成功率,行为树组合模板在较小模型的反应控制流任务上大幅恢复成功率,并且该架构可转移到运行对操作员和代理均不透明的Nav2堆栈的物理硬件上。

英文摘要

Synthesizing deployable robot behavior trees (BTs) from natural language (NL) requires grounding to ensure every generated BT references only skills a robot can actually execute. Existing LLM-based BT synthesis approaches often place this grounding responsibility on the prompt author. This makes deployment brittle when the author does not know which skills the robot can execute, how those skills are parameterized, or how the robot runtime software constrains valid BT structure. This paper proposes a contract-grounded BT synthesis architecture in which a coding agent queries a robot-side Model Context Protocol (MCP) server to retrieve an explicit contract consisting of a skill library, permitted BT operators, and optional BT composition templates, before synthesizing a BT for validation and execution. In our framework, non-expert operators issue NL commands without knowledge of robot implementation details, while a robot runtime validation gate enforces correctness before execution. We evaluate two LLMs, a closed model (Sonnet 4.6) and a smaller open-source model (Gemma4:31b), across 110 simulated tasks in PyRoboSim and 14 tasks on a physical Husarion Panther robot. Results show that contract grounding enables near-perfect BT validation and high task success, that BT composition templates substantially recover success on reactive control-flow tasks for the smaller model, and that the architecture transfers to physical hardware running a Nav2 stack opaque to both operator and agent.

CommentsIEEE RA-L Submission

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑