面向多线路公交驻站控制的完成度感知跨保真度离线到在线强化学习
Completion-Aware Cross-Fidelity Offline-to-Online Reinforcement Learning for Multi-Line Bus Holding
- Central South University(中南大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对运营车队无法探索及模拟器与目标动态差异问题,提出完成度感知的跨保真度离线到在线强化学习方法,用于多线路公交驻站控制,解决低乘客时间与不完整行程并存的失效模式。
AI中文摘要:
在运营中的公交车队上进行探索性强化学习(RL)是不切实际的,而仅从历史数据训练的策略无法获取新经验。混合离线与在线(H2O)强化学习将固定目标回放与模拟器交互相结合,但廉价的在线模拟器在转移和事件持续时间动态方面可能与目标存在差异。我们针对多线路公交驻站控制研究这一跨保真度问题,并解决一个失效模式,即较低的总乘客时间与不完整的乘客行程并存。
英文摘要:
Exploratory reinforcement learning (RL) on an operating bus fleet is impractical,while policies trained only from historical data cannot acquire new experience. Hybrid Offline-and-Online (H2O) RL combines fixed target replay with simulator interaction, but the inexpensive online simulator can differ from the target in transition and event-duration dynamics. We study this cross-fidelity problem for multi-line bus holding and address a failure mode in which lower generalized passenger time coexists with incomplete passenger journeys.