发表机构
Yonsei University(延世大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
OntoPlan提出基于本体论的场景表示与智能体框架,通过共享符号词汇对齐对象和关系,选择性检索信息并生成可执行计划,在150个任务中实现0.89成功率,令牌成本降低5.6倍。
AI 中文摘要
基于大语言模型(LLM)的机器人任务规划在开放式指令跟随方面具有前景,但在大型环境中的长时程任务上性能会下降。当空间信息通过文本传递给LLM时,模型可能无法捕获空间上下文,且令牌成本随环境规模增长。直接用LLM生成动作序列也难以满足当前世界状态和动作前置条件。我们通过一种基于本体论的场景表示来解决这一问题,该表示在共享符号词汇中对齐对象、空间、关系和状态,以支持空间推理和任务规划;并提出了OntoPlan,一个智能体框架,用于解释指令、选择性检索任务相关信息、形式化目标和约束,并生成可执行计划。在跨越五个室内环境和三种场景规模的150个一般任务中,OntoPlan实现了0.89的平均任务成功率,而最强基线为0.27,同时每个任务平均使用18.1k总令牌,比最高效的基线少约5.6倍。这些优势随着场景规模的增加而持续,而先前的方法在成功率上下降更急剧,并且在令牌成本上仍然高得多。OntoPlan还能通过询问后续问题或报告信息不足来适当响应模糊或不可行的指令,而不是承诺无效计划。代码可在https URL获取。
英文摘要
Large language model (LLM)-based robot task planning is promising for open-ended instruction following, but degrades on long-horizon tasks in large environments. When spatial information is conveyed to the LLM through text, the model can fail to capture spatial context, and token cost grows with environment size. Generating action sequences directly with an LLM also makes it difficult to satisfy the current world state and action preconditions. We address this with an ontology-grounded scene representation that aligns objects, spaces, relations, and states in a shared symbolic vocabulary for spatial reasoning and task planning, and with OntoPlan, an agentic framework that interprets instructions, selectively retrieves task-relevant information, formalizes goals and constraints, and produces executable plans. Across 150 general tasks spanning five indoor environments and three scene scales, OntoPlan achieves 0.89 average task success, compared with 0.27 for the strongest baseline, while using 18.1k total tokens per task on average, about 5.6$\times$ fewer than the most efficient baseline. These advantages persist as scene scale increases, whereas prior methods degrade more sharply in success and remain far more costly in tokens. OntoPlan also responds appropriately to ambiguous or infeasible instructions by asking follow-up questions or reporting insufficient information rather than committing to invalid plans. Code available at https://github.com/namhyeongwoo/OntoPlan.
CommentsAccepted at NeurIPS 2026. 32 pages, 7 figures