arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Sim2Signal:面向交通信号控制的仿真到现实基准

Sim2Signal: Sim-to-Real Benchmarks for Traffic Signal Control

Ferdous Al Rafi, Susrik Mukherjee, Latika Liladhar Dekate, Jennifer Yawa Lavoe, Huaiyuan Yao, Shlok Mohanty, Longchao Da, Xuesong Zhou, Hua Wei

arXiv 2609.01676首次发表:更新:

发表机构

Arizona State University(亚利桑那州立大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对交通信号控制中强化学习策略的仿真到现实差距问题,提出Sim2Signal基准并评估18种缓解方法,发现最有效方法为估计差距变化而非使策略不敏感。

AI 中文摘要

强化学习在仿真环境中实现了出色的交通信号控制性能,但在仿真器中训练出的策略一旦部署到现实世界往往会失效,这种失效被称为仿真到现实(Sim-to-Real)差距。当将强化学习应用于交通信号控制时,该差距源于感知、动作执行、交通动态和控制目标等多个来源,它们的相对影响以及现有仿真到现实缓解方法的可靠性仍未得到充分理解,且该领域缺乏用于系统测量差距和评估缓解方法的标准基准。我们提出Sim2Signal基准,该基准将仿真到现实差距分解为观测、动作、转移和奖励差距,对应于基础马尔可夫决策过程(MDP)四个组件的不匹配,并在共享协议下单独引入每个差距。我们在2个基础控制器、33个差距设置以及从5个现实位置构建的10个校准网络上评估了18种缓解方法。研究发现,直接迁移在所有四种差距来源下都会持续降低性能,但降级的严重程度无法预测缓解方法的有效性;相反,缓解方法的有效性在很大程度上取决于网络和差距设置:除动作差距外,在一种情况下有效的方法可能在另一种情况下失效。最有效的方法通常会估计差距的变化,而非通过域随机化或不变表示使策略对差距不敏感。我们的代码可在该https URL获取。

英文摘要

Reinforcement learning achieves strong traffic signal control performance in simulation, yet policies trained in simulators often fail once deployed in the real world, a failure known as the Sim-to-Real gap. When RL is applied to traffic signal control, this gap arises from several sources: sensing, action execution, traffic dynamics, and the control objective. Their relative impact and the reliability of existing Sim-to-Real mitigation methods remain insufficiently understood, and the field lacks a standard benchmark for systematically measuring the gap and evaluating mitigation methods. We present Sim2Signal, a benchmark that decomposes the Sim-to-Real gap into observation, action, transition, and reward gaps, corresponding to mismatches in the four components of the underlying MDP, and induces each gap in isolation under a shared protocol. We evaluate 18 mitigation methods on 2 base controllers, across 33 gap settings and 10 calibrated networks built from 5 real-world locations. We find that direct transfer consistently degrades performance across all four gap sources, but the severity of the degradation does not predict the effectiveness of mitigation. Instead, mitigation effectiveness depends strongly on the network and gap setting: outside the action gap, a method that helps in one case may fail in another. The most effective methods generally estimate what the gap changes, rather than make the policy insensitive through domain randomization or invariant representations. Our code is available at https://github.com/DaRL-LibSignal/Sim2Signal

Comments68 pages, 49 tables, 7 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑