arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

PlannerForge:用于自动驾驶运动规划器场景测试的LLM智能体

PlannerForge: LLM Agents for Scenario-Based Testing of Motion Planners in Autonomous Driving

Yuan Gao, Sebastian Müller, Mattia Piccinini, Marc Kaufeld, Yuchen Zhang, Finn Rasmus Schäfer, Qunying Song, Johannes Betz

arXiv 2609.08965首次发表:更新:

发表机构

Technical University of Munich; TUM School of Engineering and Design; Munich Institute of Robotics and Machine Intelligence (MIRMI); University College London(慕尼黑工业大学; 慕尼黑工业大学工程与设计学院; 慕尼黑机器人与机器智能研究所; 伦敦大学学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

PlannerForge提出统一LLM智能体框架覆盖自动驾驶场景测试全流程,扩展生成至评估并新增增强与基准阶段,在多项任务上超越现有方法,显著提升规划器成功率并降低碰撞率。

AI 中文摘要

确保自动驾驶的安全性是一个关键挑战。基于场景的测试是用于验证自动驾驶系统(ADSs)的系统性过程,但它仍然是一个碎片化的模块化流水线,其中场景生成、检索、修改、ADS执行和结果分析由相互之间几乎没有交互的独立工具完成。大型语言模型(LLM)智能体已在ADS的感知、规划和控制等子系统中展现出潜力。然而,此前没有工作使用统一的LLM智能体框架覆盖ADS的整个基于场景的测试流水线。我们提出了PlannerForge,一个LLM智能体框架,它扩展了所有基于场景的测试阶段(从场景生成到ADS评估),并增加了两个进一步的LLM增强阶段:ADS增强和ADS基准测试。我们使用10个现成的LLM,在5种提示条件下,对所有任务(生成、选择、修改、模块路由、规划器测试和增强)评估了PlannerForge。每个任务的最佳得分在0.88到1.00之间,开源20-35B后端在大多数任务上与商业API相当。诸如Qwen3.6:35B之类的开源模型在五个任务中的三个上与商业API相当。将模块端到端链接保留了83% / 78%的种子查询(商业/开源)。它在自然语言生成方面优于Scenario Factory 2.0(Finkeldei等人,2025)(200个中可执行为193个对144个),并实现了所请求的城市、道路和车辆属性的92-96%。它在排名1选择上优于BM25(Robertson和Zaragoza,2009)(92.0%对67.5%),在物理有效编辑上优于From-Words-to-Collisions(Gao等人,2025)(>=94%对31%)。在N=400时,成本调整将规划器成功率从50.4%提升至70.2%,并将碰撞率从19.0%降至8.4%,而无需特定领域的微调。

英文摘要

Ensuring the safety of autonomous driving is a critical challenge. Scenario-based testing is a systematic process used to validate Autonomous Driving Systems (ADSs), but it remains a fragmented modular pipeline in which scenario generation, retrieval, modification, ADS execution, and results analysis are performed by separate tools with little interaction. Large Language Model (LLM) agents have shown promise across ADS sub-systems such as perception, planning, and control. However, no prior work covers the whole scenario-based testing pipeline for ADSs with a unified LLM-agent framework. We present PlannerForge, an LLM-agent framework that extends all scenario-based testing stages (from Scenario Generation to ADS Assessment) and adds two further LLM-enhanced stages: ADS Enhancement and ADS Benchmarking. We evaluate PlannerForge with 10 off-the-shelf LLMs across all tasks (Generation, Selection, Modification, Module Routing, Planner Testing, and Enhancement) under 5 prompt conditions. Best-per-task scores range from 0.88 to 1.00, and open-source 20-35B backends match commercial APIs on most tasks. Open-source models such as Qwen3.6:35B match commercial APIs on three of the five tasks. Chaining the modules end-to-end retains 83% / 78% of seed queries (commercial / open). It outperforms Scenario Factory 2.0 (Finkeldei et al., 2025) on natural-language generation (193 vs. 144 executable of 200) and realises 92-96% of requested city, road and vehicle attributes. It outperforms BM25 (Robertson and Zaragoza, 2009) at rank 1 selection (92.0% vs. 67.5%) and From-Words-to-Collisions (Gao et al., 2025) on physically valid edits (>=94% vs. 31%). At N=400, cost-tuning lifts planner success from 50.4% to 70.2% and cuts collisions from 19.0% to 8.4%, without domain-specific fine-tuning.

CommentsAccepted to EMNLP 2026 (Main Conference). 35 pages including appendix

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑