发表机构
Rosenheim University of Applied Sciences(罗森海姆应用科学大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文研究基于图神经网络强化学习的课程学习在作业车间调度中的应用,通过逐步从小实例训练到大实例,显著降低训练时间并提升泛化性能,在30×30目标规模下平均最优性差距降低约8个百分点。
AI 中文摘要
作业车间调度问题是一个具有挑战性的组合优化问题,近期使用图神经网络的强化学习方法在直接从问题实例学习调度策略方面显示出潜力。然而,在大规模实例上进行训练在计算上仍然昂贵,并且跨实例规模的泛化仍然具有挑战性。本文通过将基于图神经网络的强化学习在作业车间调度问题中的课程学习与三种目标规模(20×20、25×25和30×30)下的单一规模训练进行比较,研究了课程学习。在课程设置中,策略首先在较小实例上训练,然后逐步适应较大的目标规模,使得早期阶段学到的调度行为能够支持较大实例上的学习。模型在从8×8到30×30的未见实例上使用最优性差距进行评估,同时考虑所有评估规模上的泛化性能和对目标规模的专门化性能。结果表明,课程学习始终减少墙钟训练时间,且随着目标规模的增大,收益更大。最强的优势在30×30时观察到,课程学习将所有评估规模的平均最优性差距降低了约8.1个百分点,将目标规模的平均最优性差距降低了约8.6个百分点,并节省了约50小时的训练时间。
英文摘要
The job shop scheduling problem is a challenging combinatorial optimization problem, and recent reinforcement learning approaches using graph neural networks have shown promise for learning scheduling policies directly from problem instances. However, training on large instances remains computationally expensive, and generalization across instance sizes remains challenging. This paper studies curriculum learning for graph neural network-based reinforcement learning in the job shop scheduling problem by comparing it with single-size training across three target sizes: 20 x 20, 25 x 25, and 30 x 30. In the curriculum setting, the policy is first trained on smaller instances and then progressively adapted to larger target sizes, allowing scheduling behavior learned in earlier stages to support learning on larger instances. Models are evaluated on unseen instances from 8 x 8 to 30 x 30 using the optimality gap, considering both generalization across all evaluation sizes and specialization on the target size. Results show that curriculum learning consistently reduces wall-clock training time, with larger benefits as the target size increases. The strongest advantage is observed at 30 x 30, where curriculum learning reduces the mean optimality gap across all evaluation sizes by approximately 8.1 percentage points, reduces the target-size mean optimality gap by approximately 8.6 percentage points, and saves approximately 50 hours of training time.
CommentsThis paper has been accepted for presentation at the IEEE 10th International Conference on Computational Systems and Information Technology for Sustainable Solutions (CSITSS 2026)