连接思维与行动:利用MetaTool增强的ROS框架驯服开源LLM智能体的长时程不稳定性
Bridging Thought and Action: Taming Long-Horizon Instability in Open-Source LLM Agents with a MetaTool-Enhanced ROS Framework
- Bangladesh University of Engineering and Technology (BUET)(孟加拉国工程技术大学)
- University of Massachusetts Amherst(马萨诸塞大学阿默斯特分校)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文提出MetaTool增强的ROS-Agent架构,通过强制预规划分离规划与执行,提升开源LLM智能体在长时程任务中的稳定性与效率,实验显示复杂任务完成率提升约24%。
AI中文摘要:
大型语言模型(LLMs)使得人机交互更加自然,但在智能体机器人框架中部署时,开源模型往往表现出不稳定的长时程推理和低效的动作执行。本文提出了一种基于增强型ROS-Agent的架构,以提高使用开源LLM的智能体机器人系统的任务可靠性和执行效率。所提出的系统引入了一种新颖的中间机制,称为MetaTool,它在动作执行之前强制执行结构化规划。给定自然语言命令,MetaTool引导LLM生成预期工具调用的伪代码计划,该计划存储在ROS-Agent的暂存区中,并在整个执行过程中持续存在。通过明确地将规划与执行分离,所提出的方法减少了执行循环并提高了确定性行为。该架构在具有多模态感知和运动控制能力的定制移动机器人平台上进行了验证。在真实世界交互任务上的实验结果表明,与基线框架相比,任务完成率和上下文一致性有所提高,在复杂任务上最多提升约24%。
英文摘要:
Large Language Models (LLMs) have enabled more natural human-robot interaction, but open-source models often exhibit unstable long-horizon reasoning and inefficient action execution when deployed in agentic robotic frameworks. This paper presents an enhanced ROS-Agent based architecture that improves task reliability and execution efficiency for agentic robotic systems using open-source LLMs. The proposed system introduces a novel intermediate mechanism, termed the MetaTool, which enforces structured planning prior to action execution. Given a natural-language command, the MetaTool induces the LLM to generate a pseudo-code plan of intended tool invocations, which is stored in the ROS-Agent's scratchpad and persists throughout execution. By explicitly separating planning from execution, the proposed approach reduces execution loops and improves deterministic behavior. The architecture is validated on a custom mobile robotic platform with multimodal perception and motion control capabilities. Experimental results on real-world interactive tasks demonstrate improved task completion and contextual consistency, with up to ~24% gains on complex tasks compared to the baseline framework.