AI 中文总结
HiRAD提出分层强化学习框架,通过步级时空表示、航向与速度分离及异步事件驱动流水线,解决大规模AGV连续空间路由的实时性问题,降低推理复杂度并显著减少完工时间。
AI 中文摘要
自动导引车(AGV)显著提升了仓库吞吐量,但大规模AGV车队的路由仍然具有挑战性。经典的多智能体路径规划求解器面临组合复杂度爆炸和超二次运行时间的问题,同时依赖理想化的网格或分段线性运动模型,这些模型与真实世界的运动学不匹配。最近的强化学习(RL)解决方案通过分散的智能体策略提高了灵活性,但依赖于离散化的时空表示,需要数百万个回合才能收敛,并且每一步都进行全地图观测,这导致模型庞大、收敛缓慢以及推理延迟高,违反了实时工业控制约束。为了解决这些瓶颈,我们提出了HiRAD,一个用于连续空间AGV路由的分层强化学习框架,具有实时保证:(1)一种步级时空表示,将连续运动转化为可微的强化学习问题,(2)一种分层策略,将航向选择与速度控制分离以减小动作空间,(3)一种异步事件驱动的决策流水线,将推理复杂度从O(n^2)降低到O(n),并将每步延迟最多降低71%。在随机图和两个仓库地图上,HiRAD将完工时间减少了45%至63%,并缩短了端到端运行时间。
英文摘要
Automatic Guided Vehicles (AGVs) substantially boost warehouse throughput, but routing large-scale AGV fleets remains challenging. Classical Multi-Agent Pathfinding solvers suffer from exploding combinatorial complexity and super-quadratic runtime, while relying on idealized grid or piecewise-linear motion models that mismatch real-world kinematics. Recent Reinforcement Learning (RL) solutions improve flexibility via decentralized agent policies but depend on discretized spatiotemporal representations, require millions of episodes to converge, and incur full-map observation at every step, which leads to large models, slow convergence, and high inference latency that violates real-time industrial control constraints. To address these bottlenecks, we propose HiRAD, a hierarchical RL framework for continuous-space AGV routing with real-time guarantees: (1) a step-level spatiotemporal representation that translates continuous motion into a differentiable RL problem, (2) a hierarchical strategy that splits heading choice from velocity control to reduce the action space, and (3) an asynchronous event-driven decision pipeline that lowers inference complexity from O(n^2) to O(n) and cuts per-step latency by as much as 71 percent. Across random graphs and two warehouse maps, HiRAD reduces makespan by 45 percent to 63 percent and shortens end-to-end runtime.