arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.34966cs.AI

使用拉格朗日约束的PPO与Kolmogorov-Arnold网络实现安全的温室气候控制

Safe Greenhouse Climate Control Using Lagrangian-Constrained PPO with Kolmogorov-Arnold Networks

Hangzun Liu, Yuling Fan, Fang Tian, Zhilong Bie, Zaiwen Feng, Yongliang Qiao

首次发表
浏览论文内容

中文总结 AI 辅助

提出基于拉格朗日约束的RCPO-PPO与KAN网络,将温室气候控制建模为CMDP,分离经济与安全约束,减少18.65%违规并提升2.91%利润。

中文摘要 AI 辅助

温室气候控制需要在经济效益与维持温度、湿度和CO2在作物适宜生长范围内之间取得平衡。传统的强化学习(RL)温室控制器使用固定的奖励惩罚来限制气候约束违规,然而这种启发式惩罚无法明确约束长期累积违规。调整不当的权重要么导致策略过于保守和产量降低,要么无法抑制危害光合作用并诱发作物疾病的持续气候偏差。为解决此问题,我们将温室气候调节建模为约束马尔可夫决策过程(CMDP),并使用拉格朗日安全RL框架RCPO-PPO将经济优化与累积安全约束分离,实现无需手动调整的自适应惩罚调整。为处理温室微气候与作物生长之间的强非线性、时变耦合,Kolmogorov-Arnold网络(KANs)替代多层感知机(MLPs)作为策略和价值近似器,以增强非线性表示能力。在观测中嵌入正弦周期时间特征以捕捉昼夜环境周期性。模拟使用经典冬季生菜温室模型,由40天真实天气扰动驱动。与基于惩罚的普通PPO相比,我们的方法将累积气候违规减少18.65%,并将生菜经济利润提高2.91%,同时将违规稳定保持在安全阈值附近。这种解耦的CMDP优化与基于KAN的策略表示减轻了长期气候风险并提高了种植利润,为精准温室栽培提供了一种约束感知的控制策略。

英文摘要

Greenhouse climate control balances economic return with maintaining temperature, humidity and CO2 within crop-adapted growth ranges. Conventional reinforcement learning (RL) greenhouse controllers use fixed reward penalties to limit climate constraint violations, yet such heuristic penalties cannot explicitly constrain long-term cumulative violations. Poorly tuned weights either lead to overly conservative policies and lower yields, or fail to suppress persistent climate deviations that harm photosynthesis and induce crop diseases. To address this issue, we formulate greenhouse climate regulation as a Constrained Markov Decision Process (CMDP) and use a Lagrangian safe RL framework RCPO-PPO to separate economic optimization and cumulative safety constraints, enabling adaptive penalty adjustment without manual tuning. To handle strong nonlinear, time-varying coupling between greenhouse microclimate and crop growth, Kolmogorov-Arnold Networks (KANs) replace Multi-Layer Perceptrons (MLPs) as policy and value approximators for improved nonlinear representation. Sinusoidal cyclic time features are embedded in observations to capture diurnal environmental periodicity. Simulations use a classic winter lettuce greenhouse model driven by 40-day real weather disturbances. Compared with vanilla penalty-based PPO, our method cuts cumulative climate violations by 18.65% and raises lettuce economic profit by 2.91%, keeping violations stable near the safety threshold. This decoupled CMDP optimization with KAN-based policy representation mitigates long-term climate risks and boosts planting profits, offering a constraint-aware control strategy for precision greenhouse cultivation.

发表机构

  • Huazhong Agricultural University(华中农业大学)
  • Yunnan Modern Agricultural Industry Research Institute Co., Ltd.(云南现代农业产业研究院有限公司)
  • Jilin University(吉林大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑