arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于LLM链式任务规划的设计与评估:面向通用服务机器人

Design and Evaluation of LLM Chaining-Based Task Planning for General Purpose Service Robots

Lucas Da Mota Bruno, Jiahao Sim, Yoshinobu Hagiwara

arXiv 2609.29043首次发表:更新:

发表机构

Soka University(创价大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对GPSR任务,提出LLM链式架构分离指令分类与动作生成,减少提示长度约45%,提升规划一致性,并在真实机器人上验证执行层为剩余瓶颈。

AI 中文摘要

通用服务机器人(GPSR)任务,如RoboCup@Home基准所定义,要求机器人在真实家庭环境中解释多样化的自然语言指令并生成多步动作序列。传统的单提示(SP)方法面临上下文膨胀和“中间丢失”现象,导致任务规划不可靠。我们提出了一种LLM链式架构,将指令分类和动作生成分离为两个专门阶段,将每次推理的提示长度减少约45%,同时提高规划一致性。我们使用100条随机生成的GPSR指令,在涵盖本地开源和前沿云端部署环境的三种语言模型上评估了我们的方法。结果显示,与SP相比,所有模型上的规划性能均有一致提升,本地模型上的增益高达+37个百分点。此外,在丰田人类支持机器人(HSR)上的真实机器人执行实验表明,规划成功并不保证任务完成,10个任务中有6个成功完成,执行层故障被确定为主要的剩余瓶颈。

英文摘要

General Purpose Service Robot (GPSR) tasks, as defined in the RoboCup@Home benchmark, require robots to interpret diverse natural language commands and generate multi-step action sequences in real home environments. Conventional Single Prompt (SP) approaches suffer from context bloat and the "Lost in the Middle" phenomenon, leading to unreliable task planning. We propose an LLM chaining architecture that separates instruction classification and action generation into two specialized stages, reducing per-inference prompt length by approximately 45% while improving planning consistency. We evaluate our method using 100 randomly generated GPSR commands across three language models spanning local open-source and frontier cloud deployment contexts. Results show consistent planning improvements over SP across all models, with gains of up to +37 percentage points on local models. Further, real-robot execution experiments on the Toyota Human Support Robot (HSR) reveal that planning success alone does not guarantee task completion, with 6 of 10 tasks completing successfully and execution-layer failures identified as the primary remaining bottleneck.

CommentsAccepted to IEEE GCCE 2026. 5 pages, 6 figures, 3 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑