arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

利用强化学习结合语义站点嵌入缓解公交车扎堆问题

Mitigating Bus Bunching with Reinforcement Learning Enhanced by Semantic Stop Embedding

Xin Dong, Vikash V. Gayah

arXiv 2608.10207首次发表:更新:

发表机构

The Pennsylvania State University(宾夕法尼亚州立大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出LLM辅助的语义站点嵌入强化学习方法,在仿真中较Daganzo基线显著降低公交车扎堆及乘客等待时间,实现跨线路策略复用。

AI 中文摘要

公交车扎堆会降低服务规律性并增加高频公交系统的乘客等待时间。现有基于强化学习的站点保持控制器主要依赖瞬时运营变量或特定线路的站点标识符,这些信息无法充分体现单个站点的功能和运营背景,限制了策略在不同线路间的复用。本研究引入一种由大语言模型(LLM)辅助的语义站点表示方法,用于事件驱动型公交车站点保持控制。研究中采用LLM离线将异构站点信息(包括物理属性、周边活动背景及历史运营特征)转换为固定语义嵌入,并将其融入深度Q学习控制器,无需实时进行LLM推理。实验在基于两条公交线路观测数据校准的随机仿真环境中开展,与校准后的最优Daganzo基线相比,语义控制器将车头时距变异性、扎堆事件数量及乘客等待时间分别降低了32.0%、69.2%和24.0%。特定线路的站点标识符仅用于仅考虑间距的控制器,而语义站点信息则能提升车头时距规律性、降低等待时间并减少保持操作量,在多个控制目标间实现更优权衡。跨线路实验进一步显示,零样本迁移的即时泛化能力有限,而热启动微调可加速早期学习并提升迁移策略性能;冷启动训练仍能取得最优最终性能。这些发现表明,语义状态表示可补充传统运营状态,支持相关公交线路间基于适应的策略复用。

英文摘要

Bus bunching degrades service regularity and increases passenger waiting in high-frequency transit. Existing reinforcement-learning-based holding controllers primarily rely on instantaneous operational variables or route-specific stop identifiers, which provide limited information about the functional and operational context of individual stops and constrain policy reuse across routes. This study introduces an LLM-assisted semantic stop representation for event-driven bus holding control. An LLM is used offline to transform heterogeneous stop information, including physical attributes, surrounding activity context, and historical operational characteristics, into fixed semantic embeddings that are incorporated into a deep Q-learning controller without requiring real-time LLM inference. Experiments are conducted in stochastic simulations calibrated with observed data from two bus routes. Compared with the best calibrated Daganzo baseline, the semantic controller reduces headway variability, bunching events, and passenger waiting time by 32.0%, 69.2%, and 24.0%, respectively. A route-specific stop identifier does not improve the spacing-only controller, whereas semantic stop information improves headway regularity, waiting time, and holding effort, providing a more favorable overall trade-off across control objectives. Cross-route experiments further show that zero-shot transfer provides limited immediate generalization, while warm-start fine-tuning accelerates early-stage learning and improves transferred policies; cold-start training nevertheless achieves the best final performance. These findings suggest that semantic state representations can complement conventional operational states and support adaptation-based policy reuse across related transit routes.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑