基于强化学习的工业涂装场景生产调度:采用数字模型平台
Reinforcement Learning-Based Production Scheduling in an Industry-Based Coating Scenario Using the Digital Model Playground
- Osnabrück University of Applied Sciences(奥斯纳布吕克应用科学大学)
- Kiel University(基尔大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
该研究在含序列依赖调整时间等复杂因素的工业涂装场景中,采用数字模型平台训练强化学习智能体,对比基准算法验证其调度性能,为弥合学术与工业差距提供了开源框架。
AI中文摘要:
在复杂制造环境中,生产调度面临诸多挑战,需同时考虑依赖序列的调整时间、随机干扰和交货期约束。尽管强化学习(RL)方法在研究中展现出良好效果,但多数研究依赖简化的基准流程,限制了其工业适用性。本文展示了基于RL的调度在受工业启发的涂装流程中的适用性,该流程包含依赖序列的调整时间、机器故障和可变利用率等实际复杂因素。开源数字模型平台(Digital Model Playground,DMPG)作为离散事件仿真框架,被用于对该场景建模并训练RL智能体。将深度Q网络(Deep Q-Networks)和近端策略优化(Proximal Policy Optimization)这两种标准算法与传统调度规则进行基准对比,以验证其可行性并为后续研究提供透明的测试平台。结果表明,基于RL的调度在关键性能指标上实现了均衡提升,其中PPO算法表现出最稳健的性能。本研究的主要贡献在于,通过在真实、可共享的场景中验证基于RL的调度,并为后续研究提供可复用的开源框架,从而弥合学术研究与工业实践之间的差距。
英文摘要:
Production scheduling in complex manufacturing environments is challenging when sequence-dependent setup times, stochastic disturbances, and due-date constraints must be addressed simultaneously. While reinforcement learning (RL) methods have shown promising results in research, most studies rely on simplified benchmark processes, limiting their industrial relevance. This paper demonstrates the applicability of RL-based scheduling in an industry-inspired coating process that reflects practical complexities such as sequence-dependent setup times, machine breakdowns, and variable utilization. The open-source Digital Model Playground (DMPG), a discrete event simulation framework, is used to model the scenario and to train RL agents. Two standard algorithms, Deep Q-Networks and Proximal Policy Optimization, are benchmarked against conventional dispatching rules to illustrate feasibility and to provide a transparent testbed for further research. Results indicate that RL-based scheduling achieves balanced improvements across key performance indicators, with PPO delivering the most robust performance. The main contribution of this work is to bridge the gap between academic research and industrial practice by validating RL-based scheduling in a realistic, shareable scenario and by providing a reusable open-source framework for future studies.