发表机构
Korea University of Technology and Education (KOREATECH)(韩国技术教育大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
CORAL结合五阶段课程与阶段感知奖励,在CARLA中训练多流演员-评论家策略,实现基于激光雷达的城市驾驶,长路线约束下成功率显著优于PPO基线,且可零样本跨城镇迁移。
AI 中文摘要
强化学习在自主城市驾驶领域具有应用前景,但长时程目标导向导航要求策略同时习得多种相互竞争的行为——抵达远距离目标、跟踪路线、规避障碍物、遵守交通信号,而固定目标函数未提供学习这些行为的顺序。本文提出CORAL,该方法同时推进两种调度:一是五阶段课程,逐步延长路线并收紧行为约束;二是阶段感知奖励,其组成权重随任务难度提升,将重点从任务进度转向路线跟踪、安全性、平顺性和规则合规性。策略为多流演员-评论家网络,采用近端策略优化(PPO)在CARLA中训练,使用紧凑的99维状态,该状态将极坐标激光雷达直方图与车辆遥测、自坐标系路线几何及交通规则指标配对,无需点云编码器,也无需鸟瞰图栅格化。在相同协议下与两种PPO基线对比,CORAL在完整行为约束下的最长路线的全部20次评估 episode中均达成目标,而基线的达成率仅为5%和10%;因子消融实验显示,单独任一调度都无法与两者结合的效果匹配:移除任一调度会同时降低成功率和路线完成率,同时禁用两者则成功率降至55%。在一个城镇中训练的策略,零样本迁移到7个未见城镇,在相同100-150米长度的路线上,episode成功率为68%-98%,平均横向偏差低于0.35米。
英文摘要
Reinforcement learning is promising for autonomous urban driving, but long-horizon goal-directed navigation asks a policy to acquire several competing behaviors at once--reaching a distant goal, tracking a route, avoiding obstacles, obeying signals--and a fixed objective gives no order in which to learn them. This paper presents CORAL, which advances two schedules together: a five-stage curriculum that progressively lengthens routes and tightens behavioral constraints, and a stage-aware reward whose component weights shift emphasis from mission progress toward route following, safety, smoothness, and rule compliance as the task hardens. The policy is a multi-stream actor-critic network trained with Proximal Policy Optimization (PPO) in CARLA on a compact 99-dimensional state pairing a polar LiDAR histogram with vehicle telemetry, ego-frame route geometry, and traffic-rule indicators--no point-cloud encoder, no bird's-eye-view rasterization. Against two PPO baselines under an identical protocol, CORAL reaches the goal in all twenty evaluation episodes on the longest routes under the full set of behavioral constraints, where the baselines reach 5% and 10%; a factorial ablation shows that neither schedule alone matches their combination: removing either lowers both success and route completion, and disabling both drops success to 55%. Trained in one town, the policy transfers zero-shot to seven unseen towns, succeeding in 68-98% of episodes on routes of the same 100-150 m length, with mean lateral deviation below 0.35 m.
Comments13 pages, 6 figures