arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

PORL:面向作业车间调度问题的预训练离线强化学习

PORL: Pretrained Offline Reinforcement Learning for the Job Shop Scheduling Problem

Mateo Toro Diz, Jonathan Hoss, Noah Klarmann

arXiv 2609.30948首次发表:更新:

发表机构

Rosenheim University of Applied Sciences(罗森海姆应用科学大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出PORL,一种结合在线预训练与离线微调的混合强化学习方法,用于作业车间调度,通过KL约束减少策略偏移,在分布偏移下优于独立离线强化学习,且对数据质量不敏感。

AI 中文摘要

作业车间调度问题(JSSP)是工业优化中的一个基本组合优化问题。本文提出了预训练离线强化学习(PORL),这是一种将基于仿真的在线预训练与针对生产特定数据的离线微调相结合的混合方法。通过在线交互进行强化学习能够探索通用的调度策略,但通常依赖于仿真环境,并可能面临仿真与现实的差距。相比之下,离线强化学习通过从历史数据中学习来避免与环境的直接交互,但其性能受数据集质量和覆盖度的强烈影响。PORL结合了这两种范式的优势,首先通过在线交互学习通用调度策略,随后离线地将其适应于目标分布。本文引入了一种基于KL散度的策略约束,以在微调过程中限制对预训练策略的偏离。该方法在具有分布偏移的JSSP实例以及由启发式、噪声专家和随机行为策略生成的数据集上进行了评估。结果表明,PORL始终比独立的离线强化学习和所考虑的通用调度基线获得更低的最优性差距。此外,随着数据集质量的下降,其相对于独立离线强化学习的优势增大,表明对可用离线数据的质量和覆盖度的敏感性降低。这些结果表明,预训练策略的离线适应对于直接在线探索不可行的工业调度环境是一种有前景的方法。

英文摘要

The Job Shop Scheduling Problem (JSSP) is a fundamental combinatorial optimization problem in industrial optimization. This work introduces Pretrained Offline Reinforcement Learning (PORL), a hybrid approach that combines simulation-based online pretraining with offline fine-tuning on production-specific data. Reinforcement learning through online interaction enables exploration of general scheduling strategies, but typically relies on simulation environments and may suffer from a simulation-to-reality gap. In contrast, offline RL avoids direct interaction with the environment by learning from historical data, but its performance is strongly influenced by dataset quality and coverage. PORL combines the strengths of both paradigms by first learning a general scheduling policy through online interaction and subsequently adapting it offline to a target distribution. A KL-divergence-based policy constraint is introduced to limit deviations from the pretrained policy during fine-tuning. The approach is evaluated on JSSP instances with distribution shift and datasets generated from heuristic, noisy-expert, and random behavioral policies. The results show that PORL consistently achieves lower optimality gaps than standalone offline RL and the considered general scheduling baselines. Furthermore, its advantage over standalone offline RL increases as dataset quality decreases, indicating reduced sensitivity to the quality and coverage of the available offline data. The results suggest that offline adaptation of pretrained policies is a promising approach for industrial scheduling environments where direct online exploration is impractical.

CommentsThis paper has been accepted for presentation at the IEEE 10th International Conference on Computational Systems and Information Technology for Sustainable Solutions (CSITSS 2026)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑