发表机构
University of Almería; Imperial College London(阿尔梅里亚大学; 帝国理工学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出混合离线-在线多智能体强化学习框架,利用DDPG智能体实现微藻光生物反应器中pH和溶解氧的无模型多变量控制,并在工业规模反应器中验证了其鲁棒性和适应性。
AI 中文摘要
由于活细胞系统固有的非线性和动态变异性,生物过程的有效控制尤其具有挑战性。在基于微藻的光生物反应器(PBR)中,维持稳定的pH值和溶解氧(DO)水平对于最佳生长和生产力至关重要,然而它们的强耦合性和对环境波动的敏感性使得多变量控制变得困难。本研究提出了一种新颖的混合离线-在线多智能体强化学习(MARL)框架,用于同时调节pH和DO,利用深度确定性策略梯度(DDPG)智能体实现完全数据驱动且无模型的控制解决方案。智能体使用由专家系统生成的历史数据进行训练,无需直接与环境进行实验。部署后,智能体自主运行,每天持续微调其策略以适应不断变化的过程动态并抑制快速瞬态扰动。在阿尔梅里亚大学的一个开放式工业规模PBR中进行的实验验证表明,该框架能够在现实条件下维持稳定运行。结果证实,无模型的MARL控制为复杂的生物过程环境提供了一种稳健且自适应的替代方案。
英文摘要
Effective control of bioprocesses is particularly challenging due to the intrinsic nonlinearity and dynamic variability of living-cell systems. In microalgae-based photobioreactors (PBRs), maintaining stable pH and dissolved oxygen (DO) levels is critical for optimal growth and productivity, yet their strong coupling and sensitivity to environmental fluctuations make multivariable control difficult. This study proposes a novel hybrid offline-online Multi-Agent Reinforcement Learning (MARL) framework for simultaneous pH and DO regulation, leveraging Deep Deterministic Policy Gradient (DDPG) agents to achieve a fully data-driven and model-free control solution. The agents are trained using historical data generated by an expert system, eliminating the need for direct experimentation with the environment. After deployment, the agents operate autonomously, continuously fine-tuning their policies daily to adapt to evolving process dynamics and reject fast transient disturbances. Experimental validation in an open, industrial-scale PBR at the University of Almeria demonstrated the framework's capability to maintain stable operation under realistic conditions. The results confirm that model-free MARL control provides a robust and adaptive alternative for complex bioprocess environments.
CommentsConference - IFAC WC 2026