arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

OR-Transformer:将实时决策扩展至1000个物品

OR-Transformer: Scaling Real-Time Decision-Making to 1,000 Items

Shuze Daniel Liu, David Simchi-Levi, Claire Chen, Chutong Gao, Shangtong Zhang

arXiv 2609.01933首次发表:更新:

发表机构

Massachusetts Institute of Technology; Purdue University; California Institute of Technology; University of Virginia(麻省理工学院; 普渡大学; 加州理工学院; 弗吉尼亚大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对大规模供应链补货的实时决策难题,提出OR-Transformer深度强化学习框架,其性能优于基线方法且决策时间大幅缩短,实现了千级物品规模下的实时决策。

AI 中文摘要

现代供应链运营需要在相关随机需求、异质提前期和共享固定订购成本下协调数千种异质物品的补货,产生的观测空间维度超过10^4。在该规模下,滚动时域随机混合整数线性规划(MILP)变得极其缓慢,而标准强化学习(RL)方法在高维动作空间中面临日益严峻的信用分配问题。我们提出OR-Transformer,这是一种用于随机需求下联合补货的深度强化学习框架,采用物品置换等变Transformer架构,并通过库存动态进行路径梯度训练。在最多1024个库存物品的问题规模下,OR-Transformer的性能随规模增长逐渐优于基于学习的方法和滚动时域MILP基线,且与MILP求解器相比,其在线决策时间缩短了400万倍以上,使供应链运营中的实时大规模深度强化学习成为可能。

英文摘要

Modern supply chain operations can require coordinating replenishment across thousands of heterogeneous items under correlated stochastic demand, heterogeneous lead times, and shared fixed ordering costs, yielding observation spaces exceeding $10^4$ dimensions. At this scale, rolling-horizon stochastic mixed-integer linear programs (MILPs) become prohibitively slow, while standard reinforcement learning (RL) methods face increasingly challenging credit assignment in high-dimensional action spaces. We introduce OR-Transformer, a deep reinforcement learning framework for joint replenishment under stochastic demand, with an item-permutation-equivariant Transformer architecture and pathwise-gradient training through the inventory dynamics. Across problem sizes up to 1,024 inventory items, OR-Transformer increasingly outperforms learning-based and rolling-horizon MILP baselines as scale grows. It also reduces online decision-making time by over 4 million times relative to MILP solvers, enabling real-time, large-scale deep RL in supply chain operations.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑