arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于逆模型的模型强化学习在模块化生产系统分布式优化中的应用

Distributed Optimization of Modular Production Systems using Model-based Reinforcement Learning with Inverse Models

Andreas Schwung, Steve Yuwono, Sofiene Lassoued, Dorothea Schwung

arXiv 2609.11615首次发表:更新:

发表机构

South Westphalia University of Applied Sciences; Hochschule Düsseldorf University of Applied Sciences(南威斯特法伦应用科学大学; 杜塞尔多夫应用科学大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出一种基于逆模型的模型强化学习框架,用于模块化制造系统的分布式优化,通过分离执行与状态动力学,提升训练速度与性能,尤其在离策略算法中效果显著。

AI 中文摘要

本文提出了一种用于高度灵活、模块化制造系统的数据驱动自学习控制的新方法。具体而言,我们采用了一种新颖的基于模型的强化学习框架,该框架在强化策略的训练中引入了近似逆过程模型。这种方法将执行动力学与状态空间动力学的学习分离开来,使得基于强化学习的训练仅在任务空间内进行。我们为近似逆模型提出了一种轻量级前馈架构,并将其集成到标准强化学习算法的策略网络中。我们将该方法应用于一个具有异构生产模块的实验室模块化生产测试平台。结果凸显了模块化制造单元在性能和训练速度方面的效率提升,尤其是在离策略算法中。

英文摘要

This paper presents a novel approach for data-driven self-learning control of highly flexible, modular manufacturing systems. Specifically, we employ a novel framework for model-based reinforcement learning which introduces approximate inverse process models within the training of reinforcement policies. This approach disentangles the learning of actuation dynamics and the dynamics in state space, resulting in RL-based training solely within the task space. We propose a lightweight feedforward architecture for approximate inverse models and integrate them within the policy network of standard RL algorithms. We apply the approach to a laboratory modular production testbed with heterogeneous production modules. The results underline the efficiency improvements for modular manufacturing units in terms of both performance and training speed, particularly for off-policy algorithms.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑