发表机构
Universidad Politécnica de Cartagena(卡塔赫纳理工大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对电动配送车队班中充电协调难题,提出基于纯本地控制的学习智能体策略,在二十城模拟中零样本部署达98.6%按时完成率,接近Oracle性能,显著优于贪婪规则和阈值启发式方法。
AI 中文摘要
在电动配送车队中,班中充电并非易事:每辆车必须决定何时、何地以及充电多少,以便在电池高于安全下限的情况下按时完成任务。这些选择是相互耦合的:当太多车辆选择同一站点时,就会形成排队。先前的工作通过中央调度、预计算时间表或预订来解决这种耦合,而这些机制是充电基础设施很少支持的。相反,我们使用一组在纯本地控制下的学习智能体:每辆车运行相同的策略,仅根据其时间预算和广播的站点占用率自行决策,从而在没有中央控制或消息传递的情况下实现涌现协调。我们在真实OpenStreetMap网络的二十个城市的模拟中验证了这一范式,每个城市都有一个由全知Oracle校准的冻结场景(99.5%的班次按时完成),而一个简单的贪婪规则(低电量时选择最近站点)仅完成73%。使用神经进化(NEAT)和策略梯度(PPO)在四个城市上训练的智能体,零样本部署到全部二十个城市(其中十六个在训练中从未见过),分别完成了96.8%和98.6%的班次,其中策略梯度控制器在需求或车辆特性偏离训练范围时表现出更强的鲁棒性。相比之下,仅读取车辆紧急程度的调整阈值启发式方法在拥挤城市中表现不佳(约80%)。通过训练,这些学习智能体重新发现了部分充电和短时机会充电,并绕开繁忙站点,将每次会话的排队等待时间从约45分钟缩短至2分钟以下。总之,这种协调范式在本地紧急程度与公共占用率之间取得平衡,以最小的实施成本达到接近Oracle的性能。
英文摘要
In electric delivery fleets, mid-shift charging is non-trivial: each vehicle must decide when, where and how much to charge to finish on time with battery above a safety floor. The choices are coupled: queues build where too many vehicles pick the same station. Prior work resolves this coupling with central dispatching, precomputed schedules or reservations, machinery that charging infrastructure rarely supports. Instead, we use a family of learning agents under purely local control: every vehicle runs the same policy, deciding alone from its time budgets and broadcast station occupancies, leading to emergent coordination without central control or messaging. We validate this paradigm in simulation on real OpenStreetMap networks of twenty cities, each with a frozen scenario calibrated by an omniscient Oracle (99.5% of shifts completed on time), whereas a naive greedy rule (nearest station on low battery) completes just 73%. Agents trained with neuroevolution (NEAT) and policy gradients (PPO) on four cities and deployed zero-shot across all twenty, sixteen never seen in training, complete 96.8% and 98.6% of shifts, with the policy-gradient controllers proving more robust when demand or vehicle characteristics drift beyond the trained regime. In contrast, tuned threshold heuristics that read vehicle urgency alone fall short in contended cities (~80%). Through training, these learning agents rediscover partial charging and short opportunistic sessions, and route around busy stations, cutting per-session queue waits from about 45 minutes to under 2. In summary, this coordination paradigm balances local urgency against public occupancy, reaching near-Oracle performance at minimal implementation cost.
Comments42 pages, 11 figures. Submitted to Transportation Research Part C