arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

增强PID控制与深度强化学习:面向工业基准的混合方法

Augmenting PID Control with Deep Reinforcement Learning: A Hybrid Approach to the Industrial Benchmark

Zhengyang, Gu, Joseph E. Hernandez, John Burtenshaw, Sean Scott, Thomas Cook, Chris Couch

arXiv 2609.22584首次发表:更新:

发表机构

Liveline Technologies(莱夫兰科技)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对工业过程复杂性,提出混合PID-RL控制器,利用TD3智能体发现最优参数并馈入PID,兼顾性能、效率与可靠性。

AI 中文摘要

随着工业过程日益复杂,传统的比例-积分-微分(PID)控制器往往难以应对其非线性、多输入动态特性。我们提出使用先进的深度强化学习(DRL)来证明其在复杂环境中的优势。为此,我们依托工业基准(IB)。IB是一个逼真的仿真环境,用于测试DRL算法应对工业应用关键挑战的能力:高维状态空间、延迟效应以及相互冲突的多准则目标。该测试平台凸显了DRL的核心权衡:虽然其最终策略往往不稳定,但其独特优势在于能够在简单控制器失效的多维空间中自主发现最优的、非显而易见的策略。在本文中,我们提出一种新颖的混合PID-RL控制器,该控制器利用DRL的发现能力,同时确保可靠性。在开发多目标奖励函数使DRL可行之后,我们使用双延迟深度确定性策略梯度(TD3)智能体作为发现工具,为IB的“增益”和“偏移”参数寻找最优的、非显而易见的设置。通过将这些发现的参数输入到简单且经过调优的PID控制器中,我们的混合模型成功结合了所有三个特性:它实现了最优DRL智能体的性能与效率,同时具备经典控制器的可靠性。这项工作展示了一种实用的方法论,即使用DRL来增强而非取代可信赖的工业控制系统。

英文摘要

As industrial processes grow in complexity, traditional Proportional-Integral-Derivative (PID) controllers are often insufficient for handling their non-linear, multi-input dynamics. We propose using advanced Deep Reinforcement Learning (DRL) to prove its advantages in these complex environments. To do this, we rely on the Industrial Benchmark (IB). The IB is a realistic simulation that tests DRL algorithms against the key challenges of industrial applications: high-dimensional state spaces, delayed effects, and conflicting multi-criterial objectives. This testbed highlights DRL's core trade-off: while its final policies can often be unstable, its unique strength is the ability to autonomously discover optimal, non-obvious policies in multi-dimensional spaces where simple controllers fail. In this paper, we propose a novel hybrid PID-RL controller that leverages DRL's discovery capability while ensuring Reliability. After developing a multi-objective reward function to make DRL viable, we use a twin-delayed deep deterministic (TD3) agent as a discovery tool to find the optimal, non-obvious settings for the IB's 'Gain' and 'Shift' parameters. By feeding these discovered parameters to a simple, tuned PID controller, our hybrid model successfully combines all three characteristics: it achieves the optimal Performance and Efficiency of the best DRL agent with the Reliability of a classical controller. This work demonstrates a practical methodology for using DRL to augment, rather than replace, trusted industrial control systems.

CommentsPublished in: 2026 7th International Conference on Artificial Intelligence, Robotics, and Control (AIRC)

DOI:10.1109/AIRC69745.2026.11631341

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑