arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

用于运行时可调公交信号优先的偏好条件多目标强化学习

Preference-Conditioned Multi-Objective Reinforcement Learning for Runtime-Tunable Transit Signal Priority

Philip-Roman Adam, Stefanie Schmidtner

arXiv 2607.18286首次发表:更新:

AI 中文总结

研究公交信号优先中平衡多目标的问题,提出偏好条件TSP控制器,能在运行时通过偏好参数调整权衡公交与整体交通延误,无需重训练。通过扩展场景生成实现,实验表明该控制器优于基线,保持约束可行性,且能揭示不同偏好下的非公交外部性情况。

AI 中文摘要

公交信号优先(TSP)需要平衡相互竞争的目标:减少公交延误,同时限制对非公交交通的不利影响,并避免部分车辆的极端等待。现有TSP的强化学习方法通常编码公交感知特征,但优化固定奖励或固定标量化,当机构优先级随时间或干扰条件变化时,限制了操作灵活性。我们提出了一种偏好条件TSP控制器,$\pi(a \mid s,w)$,它在最小/最大绿灯和转换可行性约束下选择下一个信号相位,并可通过偏好参数$w$在运行时进行调整,以权衡公交优先级强调与整体交通延误,无需重新训练。我们通过引入约束信号控制/TSP包装器在IntersectionZoo之上实现了这一点,并通过公交流行度增强和基于时间表的公交插入扩展了场景生成,以解决训练期间稀疏的公交优先事件。与固定时间控制、基于规则的TSP覆盖和固定权重PPO专家的实验表明,单个学习的条件策略在运行时偏好上跨越了一个平滑的经验权衡前沿,优于固定时间和基于规则的基线,并保持约束可行性,而尾部延误诊断表明,对于适度的偏好设置,非公交外部性仍然有限,但在高公交优先级权重下可能会大幅增加。这项工作的源代码可在这个https URL上获得。

英文摘要

Transit signal priority (TSP) requires balancing competing objectives: reducing bus delay while limiting adverse impacts on non-bus traffic and avoiding extreme waits for a subset of vehicles. Existing reinforcement-learning (RL) approaches to TSP typically encode transit-aware features (e.g., occupancy and schedule deviation) but optimize a fixed reward or fixed scalarization, which limits operational flexibility when agency priorities change across time-of-day or disruption conditions. We present a preference-conditioned TSP controller, $π(a \mid s,w)$, that selects the next signal phase under minimum/maximum green and transition-feasibility constraints and can be tuned at runtime via a preference parameter $w$ to trade off bus-priority emphasis against overall traffic delay without retraining. We implement this on top of IntersectionZoo by introducing a constrained signal-control/TSP wrapper, and we extend scenario generation with bus-prevalence augmentation and timetable-based bus insertion to address sparse transit-priority events during training. Experiments against fixed-time control, a rule-based TSP overlay, and fixed-weight PPO specialists show that a single learned conditioned policy spans a smooth empirical trade-off frontier across runtime preferences, outperforms fixed-time and rule-based baselines, and maintains constraint feasibility, while tail-delay diagnostics reveal that non-bus externalities remain limited for moderate preference settings but can increase substantially under high bus-priority weights. The source code of this work is available at https://github.com/urbanAIthi/morl-tsp.

Comments8 pages, 3 figures, 4 tables. Accepted at the 29th IEEE International Conference on Intelligent Transportation Systems (ITSC 2026). Code: https://github.com/urbanAIthi/morl-tsp

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑