arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

REARL:一种融合真实交通数据与大语言模型的闭环自动驾驶仿真增强框架

REARL: A Closed-loop Autonomous Driving Simulation Enhancement Framework with Real Traffic Data and Large Language Models

Xiaojun Bi, Jun Jiang, Yiwen Sun, Quanyi Ou, Yizhi Ma, Ke Cheng, Mingjie Bi, Yexin Li, Bowen Du

arXiv 2609.19903首次发表:更新:

发表机构

Minzu University of China; Peking University; BIGAI; Beihang University(中央民族大学; 北京大学; 北京通用人工智能研究院; 北京航空航天大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

REARL提出闭环仿真增强框架,融合真实交通数据与大语言模型,通过聚类场景和滑动窗口检测器动态调整车辆决策,在HighD高速场景中显著降低分布偏差,提升仿真真实性。

AI 中文摘要

准确的仿真对于自动驾驶开发至关重要,然而捕捉真实世界的交通复杂性仍然具有挑战性。依赖预定义规则或静态数据回放的现有仿真器难以应对动态交通。CRITICAL使用真实交通数据和大语言模型(LLM)来调整初始仿真配置,但随着推演(rollout)的进行,仿真分布仍与真实交通存在偏差。我们提出REARL,一种将真实交通数据与LLM相结合的闭环仿真增强框架。真实交通数据被聚类,每个聚类中心被用作代表性场景,为LLM提供典型的真实世界交通模式。随后,一个定时滑动窗口检测器监测车辆速度分布和车辆对之间平均间距的差异。如果某个指标超过阈值,LLM调整车辆决策;否则保留现有控制器。LLM还从交通快照中选择匹配的真实车辆,并参考该真实动作来调节仿真车辆。在受控的HighD高速公路场景中,与CRITICAL基线和基于PPO的学习基线相比,REARL将速度分布的Hellinger距离降至0.3067,平均间距的MAPE降至0.8371,同时实现了22.8575的时间车头时距(THW)和0.0708的变道率。

英文摘要

Accurate simulation is crucial for autonomous driving development, yet capturing real-world traffic complexity remains challenging. Existing simulators that rely on predefined rules or static data playback struggle with dynamic traffic. CRITICAL uses real traffic data and a large language model (LLM) to adjust the initial simulation configuration, but the simulated distribution still diverges from real traffic as the rollout evolves. We propose REARL, a closed-loop simulation enhancement framework that integrates real traffic data with LLMs. Real traffic data are clustered, and each cluster center is used as a representative scenario that provides typical real-world traffic patterns for the LLM. A timed sliding-window detector then monitors discrepancies in vehicle speed distribution and mean spacing between pairs of vehicles. If a metric exceeds a threshold, the LLM adjusts vehicle decision-making; otherwise the existing controller is kept. The LLM also selects a matching real vehicle from a traffic snapshot and modulates the simulated vehicle with reference to that real action. In a controlled HighD highway setting, compared with the CRITICAL baseline and a PPO-based learning baseline, REARL reduces the Hellinger distance for speed distributions to 0.3067 and the MAPE for mean spacing to 0.8371, while achieving a time headway (THW) of 22.8575 and a lane change rate of 0.0708.

Comments14 pages, 8 figures, 3 tables. Corresponding author: Yiwen Sun. This work was supported by the National Natural Science Foundation of China (Grant No. 62503015)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑