arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

物理智能体AI:一种用于协调机器人团队与大语言模型的架构

Physical Agentic AI: An Architecture for Orchestrating a Robot Crew with LLMs

Xinyuan Liu, Eren Sadikoglu, Riana Chatterjee, Ransalu Senanayake

arXiv 2608.22657首次发表:更新:

发表机构

School of Computing and Augmented Intelligence, Arizona State University(亚利桑那州立大学计算与增强智能学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出Physical Agentic AI框架,通过语义规划与执行间的显式架构接口协调机器人团队,在两类任务中验证其可提升技能落地率、减少故障动作,为物理机器人团队提供可靠协调方案。

AI 中文摘要

智能体AI框架可解读开放式任务目标并将其分解为多步骤计划。关于具身特定能力、物理前提条件及跨机器人协调的丰富信息提升了落地性,但无法消除不可行、时机不当或不安全的物理动作。因此,物理机器人团队需要在语义规划与执行之间设置显式架构接口,每个规划动作在执行前都需依据机器人能力、系统状态及工作流约束进行验证。本文提出Physical Agentic AI,一种基于技能的机器人智能体协调框架,其中每个机器人暴露一个类型化的可执行技能库,而基础模型规划器将任务分解为多个阶段并为每个阶段分配机器人-技能对。机器人协调层将技能库、机器人状态、命名位置及工作流契约暴露给非执行任务规划器,同时确定性机器人协调器每次验证并授权一个技能。我们在无人机-无人地面车辆(UGV)搜索与调度任务(所有条件下的每个任务均在Gazebo中实时执行)上,以及在人形机器人-四足机器人运输任务(使用硬件等效技能接口,加上在Unitree G1与Go2上的两次物理试验)上进行评估。通过独立改变规划器知识与运行时约束,我们发现检索将技能落地率从51%提升至96%,但仍导致知情规划器调度23-29%的故障步骤;每次调度约束将错误调度降至0%且无错误阻塞,保留计划消融实验确认是约束门而非计划变化导致该结果。实时执行使物理层面产生差异:无约束时,所有8个注入故障均越过协调边界,6个导致机器人运动;有约束时,所有8个故障在运动前被拒绝。

英文摘要

Agentic AI frameworks interpret open-ended task goals and decompose them into multi-step plans. Richer information about embodiment-specific capabilities, physical preconditions, and cross-robot coordination improves grounding, but does not eliminate infeasible, mistimed, or unsafe physical actions. Physical robot crews therefore require an explicit architectural interface between semantic planning and execution, where every planned action is verified against robot capabilities, system state, and workflow constraints before actuation. This paper introduces Physical Agentic AI, a framework for skill-grounded robot agent orchestration, in which each robot exposes a typed library of executable skills while a foundation model planner decomposes a task into phases and assigns each phase to a robot-skill pair. A Robot Orchestration layer exposes the skill library, robot state, named locations, and workflow contracts to a non-actuating Mission Planner, while a deterministic Robot Orchestrator validates and authorizes one skill at a time. We evaluate on a drone-UGV search-and-dispatch mission, where every mission in every condition is executed live in Gazebo, and on a humanoid-quadruped transportation task using hardware-equivalent skill interfaces plus two physical trials on a Unitree G1 and Go2. Varying planner knowledge and runtime enforcement independently, we find that retrieval raises skill grounding from 51% to 96% yet leaves informed planners dispatching 23-29% of faulted steps. Per-dispatch enforcement reduces false dispatch to 0% with no false blocks, and a held-plan ablation confirms that the gate, not plan variation, is responsible. Live execution makes the difference physical: without enforcement all eight injected faults crossed the orchestration boundary and six produced robot motion; with enforcement all eight were refused before motion.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑