GAPL:面向基于大语言模型的轨迹规划的接地动作效果策略学习
GAPL: Grounded Action-effect Policy Learning for LLM-Based Trajectory Planning
浏览论文内容
中文总结 AI 辅助
针对LLM用于自动驾驶轨迹规划时存在的幻觉、接地不足等问题,提出GAPL框架,整合三类模块并经实验验证其性能优于基线方法
中文摘要 AI 辅助
自动驾驶的轨迹规划既需要高层推理能力,也需要精准的底层控制能力。大语言模型(LLM)具备语义丰富的规划能力,但其应用受限于推理幻觉、环境动态性接地不足以及控制数值精度有限等问题。本文提出GAPL(接地动作效果策略学习),这是一个将基于LLM的效果估计、基于仿真的效果接地和策略优化整合为闭环系统的统一框架。GAPL包含三个模块:(1)用于结构化多维动作效果估计的基于LLM的效果评估器;(2)基于仿真的效果接地器,通过仿真器回滚预测与动态一致的效果;(3)效果感知决策器,该决策器通过蒸馏器将LLM的效果估计与仿真相接地,以指导基于近端策略优化(PPO)的策略学习。在四个Highway-env场景上开展的实验表明,GAPL的性能始终优于基线方法,其碰撞率、平均位移误差(ADE)和最终位移误差(FDE)分别实现了{0.76, 0.86, 2.00}的平均降低,平均奖励增益为1.44。
英文摘要
Trajectory planning for autonomous driving requires both high-level reasoning and precise low-level control. Large Language Models (LLMs) offer semantic-rich planning capabilities, however, their application is limited by hallucinated reasoning, poor grounding in environment dynamics, and limited numerical precision in control. We propose GAPL (Grounded Action-effect Policy Learning), a unified framework that integrates LLM-based effect estimation, simulation-based effect grounding, and policy optimization into a closed-loop system. GAPL consists of three modules: (1) an LLM-based Effect Evaluator for structured multi-dimensional action-effect estimation; (2) a Simulation-based Effect Grounder that predicts dynamics-consistent effects from simulator rollouts; and (3) an Effect-Aware Decision Maker that grounds LLM effect estimates against simulation via a distiller to guide Proximal Policy Optimization (PPO)-based policy learning. Experiments on four Highway-env scenarios demonstrate that GAPL consistently outperforms baselines, achieving average reductions of {0.76, 0.86, 2.00} in collision rate, average displacement error (ADE), and final displacement error (FDE), and an average reward gain of 1.44.
发表机构
- University of Oslo(奥斯陆大学)
- Aalborg University(奥尔堡大学)
- University of Sydney(悉尼大学)
- Nanjing University of Information Science and Technology(南京信息工程大学)
- Simula Research Laboratory(西穆拉研究院)
机构由 AI 辅助整理,请以论文原文为准。