arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AeroEval:AI生成无人机任务的分阶段程序与执行验证

AeroEval: Staged Program and Execution Validation for AI-Generated Drone Missions

Kautuk Astu, Naina Rabha, Yogesh Simmhan

arXiv 2610.09764首次发表:更新:

发表机构

Indian Institute of Science(印度科学研究所)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

AeroEval提出分阶段程序与执行验证中间件,结合确定性分析和LLM智能体,将无人机任务成功率从55%提升至95%,有效检测并修正导航与分析任务中的多种故障。

AI 中文摘要

大型语言模型(LLMs)能够根据自然语言任务描述生成无人机程序,但语法上有效的程序仍可能违反用户意图、环境约束和任务级行为。这一问题在网络物理应用中尤为突出,因为其正确性取决于生成代码、移动感知、环境几何、事件驱动分析和物理执行之间的交互。现有的无人机代码生成系统主要依赖提示护栏或模拟器结果,且故障定位能力有限。我们提出了AeroEval,一种用于AI生成无人机任务分阶段验证的智能体辅助中间件。AeroEval将确定性程序分析与基于上下文的LLM智能体相结合。它首先验证程序语法、平台API使用和任务意图,然后利用执行轨迹、任务需求和环境上下文评估实际行为。每个阶段都会返回结构化故障信息,以供迭代重新生成。在我们使用20个导航任务和五种分析型任务类型在AirSim和Gazebo模拟器上的评估中,AeroEval将导航成功率从55%提升至95%。在分阶段消融研究中,仅使用代码和轨迹验证器分别达到44%和56%的平均运行级成功率,而完整的AeroEval流程达到88%;各阶段能检测程序结构、API使用、任务意图、避障、高度、覆盖范围和事件驱动转换中的互补性故障,且引导式重新生成能修正这些故障。在主要分析任务中,AeroEval将聚合运行级成功率从一次性AeroGen的34%提升至重新生成预算内的88%。这些结果表明,在评估环境中,将程序级与执行接地智能体验证相结合对AI生成的无人机应用具有显著益处。

英文摘要

Large Language Models (LLMs) can generate drone programs from natural-language mission descriptions, but syntactically valid programs may still violate user intent, environmental constraints, and mission-level behavior. This problem is pronounced in cyber-physical applications, where correctness depends on the interaction among generated code, mobile sensing, environmental geometry, event-driven analytics, and physical execution. Existing drone code-generation systems primarily use prompt guardrails or simulator outcomes and provide limited failure localization. We present AeroEval, an agent-assisted middleware for staged validation of AI-generated drone missions. AeroEval combines deterministic program analysis with context-grounded LLM agents. It first validates program syntax, platform API usage, and mission intent, and then evaluates the realized behavior using execution trajectories, mission requirements, and environmental context. Each stage returns structured failure information for iterative regeneration. In our evaluation using 20 navigation tasks and five analytical mission types over AirSim and Gazebo simulators, AeroEval improves navigation success from 55% to 95%. In a stagewise ablation study, our Code and Trajectory Validators by themselves achieve mean run-level success rates of 44% and 56%, respectively, while the full AeroEval pipeline achieves 88%; the stages detect complementary failures in program structure, API usage, mission intent, obstacle avoidance, altitude, coverage, and event-driven transitions and the guided regeneration corrects for them. Across the main analytics missions, AeroEval increases aggregate run-level success from 34% for one-shot AeroGen to 88% within the regeneration budget. These results demonstrate the benefit of combining program-level and execution-grounded agentic validation for AI-generated drone applications in the evaluated environment.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑