发表机构
The University of Hong Kong; University of California, Berkeley(香港大学; 加利福尼亚大学伯克利分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对大规模 OHT 系统路由问题,提出轨迹初始化的神经双 Q 路由算法,在多组集群规模-到达率设置中显著降低平均完成时间,提升已完成任务数量并缩短尾部完成时间。
AI 中文摘要
大规模工业机器人集群共享受限的物理基础设施,导致车辆行驶时间取决于安全间隔、交叉口通行权、下游阻塞以及站点竞争。本文以半导体工厂中典型的天花板式物料搬运系统—— overhead hoist transport(OHT)系统为研究对象,针对静态最短路径路由无法考虑这些时变交通成本、表格型 Q 路由虽能在线自适应但因独立学习各目的地-节点-动作值而导致稀疏访问路由场景间信息共享受限、启动行为受不准确值估计影响的问题,提出 Neural Double Q-routing 算法,该算法用共享状态-动作值网络替代按目的地索引的表格;通过对混合模拟器生成的路由轨迹进行回报回归实现网络的热启动,再利用双 Q 更新、局部拥塞校正以及按事件分层的结构化回放进行在线优化。在 100、150、200 台 OHT 的 9 组匹配的集群规模-到达率设置中,所提框架相对表格型 Double Q-routing 将平均完成时间降低了 0.8%至 8.8%;在 150 和 200 台 OHT 的 6 组设置中,其平均完成时间为所有对比方法中最低,而 Dijkstra 算法在 100 台 OHT 的 3 组设置中仍表现最佳;在 9 组设置中的 8 组,已完成任务数量与表格型 Double Q-routing 的差值在 1%以内,且 8 组设置的第 95 百分位完成时间有所降低;在 2 组匹配的启动场景中,离线初始化使已完成任务数量最多提升 23%,尾部完成时间最多降低 15%。
英文摘要
Large-scale industrial robot fleets share constrained physical infrastructure, making vehicle travel times dependent on safety separation, intersection access, downstream blocking, and station contention. We study this problem in overhead hoist transport (OHT) systems, a representative ceiling-mounted material-handling system used in semiconductor fabs. Static shortest-path routing cannot account for these time-varying traffic costs, whereas tabular Q-routing adapts online but learns each destination--node--action value independently, limiting information sharing across sparsely visited routing contexts and making startup behavior sensitive to inaccurate value estimates. We propose Neural Double Q-routing, which replaces destination-indexed tables with a shared state--action value network. The network is warm-started through return-to-go regression on mixed simulator-generated routing trajectories and then refined online using Double-Q updates, local congestion correction, and event-stratified structured replay. Across nine matched fleet-size--arrival-rate settings with 100, 150, and 200 OHTs, the proposed framework reduces mean completion time relative to tabular Double Q-routing by $0.8\%$--$8.8\%$. It achieves the lowest mean completion time among all compared methods in the six 150- and 200-OHT settings, whereas Dijkstra remains best in the three 100-OHT settings. Completed-task counts remain within $1\%$ of tabular Double Q-routing in eight of nine settings, and 95th-percentile completion time decreases in eight settings. In two matched startup scenarios, offline initialization increases the number of completed tasks by up to $23\%$ and reduces tail completion time by up to $15\%$.