arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.25397cs.LGcs.AImath.OC

一维装箱问题的物品兼容图深度强化学习

Deep Reinforcement Learning on Item-Compatibility Graphs for One-Dimensional Bin Packing

  • Bahçeşehir University(巴哈切希尔大学)

机构由 AI 辅助整理,请以论文原文为准。

M. Aslı Aydın

AI总结:

提出基于物品兼容图的深度强化学习框架,通过图神经网络策略和随机束搜索,零样本泛化求解一维装箱,降低构造式启发式的平均最优性差距至2.31%。

AI中文摘要:

一维装箱问题(1D-BPP)是一个经典的NP难组合优化问题,其应用范围涵盖物流、制造和云资源管理等领域。尽管深度强化学习(DRL)已成为数据驱动优化的竞争性范式,但大多数学习型装箱方法针对2D和3D变体,而针对1D-BPP的智能学习求解器仍然稀缺。本文提出了一种新颖的端到端、尺寸无关的图强化学习框架用于1D-BPP。我们将装箱过程建模为物品兼容图上的马尔可夫决策过程,该图作为一种结构化知识表示,其中每个动作合并两个可匹配的部分装箱。图神经网络演员-评论家策略从该表示中提取关系特征,并通过强化学习训练,由随机束搜索解码,使得单个训练模型能够零样本泛化到任意规模的实例。我们针对图编码器、DRL算法、奖励函数、训练分布和超参数进行了系统的实证研究。在完整BPPLIB基准上零样本评估,与构造式启发式、分组遗传算法和近期学习型方法相比,我们的数据驱动策略将构造式启发式的平均最优性差距从2.66%降至2.31%,在结构化实例上提升最大。与在同一基准上评估的学习型基线相比,在九个族中的大多数上实现了更低的差距,并且在实例分布上更加稳定。在最难的基准族上,它优于依赖列生成和整数规划的先进学习型求解器,而完全不使用求解器。分组遗传算法总体上仍保持领先,我们分析了残余差距出现的位置和原因。

英文摘要:

The one-dimensional bin packing problem (1D-BPP) is a classical NP-hard combinatorial optimization problem with applications ranging from logistics and manufacturing to cloud resource management. Although deep reinforcement learning (DRL) has become a competitive paradigm for data-driven optimization, most learned packing methods target 2D and 3D variants, and intelligent learned solvers for 1D-BPP remain scarce. In this paper, we present a novel end-to-end, size-agnostic graph reinforcement learning framework for 1D-BPP. We formulate the packing process as a Markov decision process on an item-compatibility graph, serving as a structural knowledge representation in which every action merges two partial bins that fit together. A graph neural network actor-critic policy extracts relational features from this representation and is trained through reinforcement learning and decoded by stochastic beam search, enabling a single trained model to generalize zero-shot to instances of any size. We conduct a systematic empirical study across graph encoders, DRL algorithms, reward functions, training distributions, and hyperparameters. Evaluated zero-shot on the full BPPLIB benchmark against a constructive heuristic, a grouping genetic algorithm, and recent learned methods, our data-driven policy lowers the mean optimality gap of the constructive heuristic from 2.66\% to 2.31\%, with the largest gains on structured instances. Against learned baselines evaluated on the same benchmark, it attains a lower gap on most of the nine families and is far more stable across instance distributions. On the hardest benchmark family, it outperforms a state-of-the-art learned solver that relies on column generation and integer programming, while using no solver at all. A grouping genetic algorithm remains ahead overall, and we analyze where and why the residual gap arises.

补充信息

↑