arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.23695cs.AI

PhysAI-Bench:面向自主无人机为中心的物理AI中基于LLM的智能体决策基准

PhysAI-Bench: A Benchmark for LLM-Based Agentic Decision-Making in Autonomous UAV-Centric Physical AI

Mohamed Amine Ferrag, Merouane Debbah, Abderrahmane Lakas, Manu Perumkunnil, Norbert Tihanyi

首次发表
浏览论文内容

中文总结 AI 辅助

提出PhysAI-Bench基准,含10,178个无人机任务决策实例,评估29个基础模型,GPT-5.3准确率最高52%,揭示物理AI智能体决策仍是开放挑战。

中文摘要 AI 辅助

物理AI的最新进展加速了基础模型在自主系统(如无人机)中的应用,这些系统必须在动态环境中感知、推理、规划并行动。现有基准评估物理感知、直觉物理、具身导航和协作推理,但很少评估实现可靠自主性所需的智能体决策能力。我们引入了PhysAI-Bench,一个用于评估该能力的基准。它包含10,178个标准化决策实例,这些实例从自主无人机任务对话轨迹中自动提取。每个实例保留任务上下文、时间依赖关系、物理约束、模型上下文协议(MCP)工具调用、智能体间(A2A)交互、传感器观测以及AI原生6G网络条件,包括延迟、丢包率、吞吐量、边缘负载和网络切片。我们仅暴露每个决策之前的信息,防止未来事件泄漏并近似在线决策。我们使用两阶段协议评估29个基础模型。我们从零样本、三样本和五样本提示的12种组合及四种温度中,在35个实例的人工验证开发集上进行三次运行,选择模型特定配置。然后冻结每个选定配置,并在固定的、情节不相交的500个实例集上进行三次运行评估。GPT-5.3达到最高准确率(52.00%),其次是GPT-5.2(49.40%)和Grok 4.5(49.07%)。少样本提示通常提高性能,而温度影响有限。结果表明,物理AI中可靠的智能体决策仍是一个开放挑战。数据集可在该https URL获取。

英文摘要

Recent advances in Physical AI have accelerated the use of foundation models in autonomous systems such as unmanned aerial vehicles (UAVs), which must perceive, reason, plan, and act in dynamic environments. Existing benchmarks assess physical perception, intuitive physics, embodied navigation, and collaborative reasoning, but rarely evaluate the agentic decision-making required for reliable autonomy. We introduce \textit{PhysAI-Bench}, a benchmark for evaluating this capability. It contains 10,178 standardized decision instances automatically extracted from conversational traces of autonomous UAV missions. Each instance preserves mission context, temporal dependencies, physical constraints, Model Context Protocol (MCP) tool calls, Agent-to-Agent (A2A) interactions, sensor observations, and AI-native 6G network conditions, including latency, packet loss, throughput, edge load, and network slicing. We expose only information preceding each decision, preventing future-event leakage and approximating online decision-making. We evaluate 29 foundation models using a two-stage protocol. We select model-specific configurations from 12 combinations of zero-, three-, and five-shot prompting and four temperatures, tested in three runs on a 35-instance, human-verified development set. We then freeze each selected configuration and evaluate it in three runs on a fixed, episode-disjoint set of 500 instances. GPT-5.3 achieves the highest accuracy (52.00%), followed by GPT-5.2 (49.40%) and Grok~4.5 (49.07%). Few-shot prompting generally improves performance, while temperature has limited influence. The results demonstrate that reliable agentic decision-making in Physical AI remains an open challenge. The dataset is available at https://github.com/maferrag/physai-bench

发表机构

  • United Arab Emirates University(阿联酋大学)
  • Khalifa University(哈利法大学)
  • Interuniversity Microelectronics Centre (IMEC)(校际微电子中心)
  • Technology Innovation Institute(技术创新研究所)

机构由 AI 辅助整理,请以论文原文为准。

↑