arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

可微混合动作神经反馈控制用于区域供热网络

Differentiable Hybrid-Action Neural Feedback Control for District Heating Networks

Nicolas Kirsch, Corrado Sgadari, Alessio La Bella, Giancarlo Ferrari-Trecate

arXiv 2610.01822首次发表:更新:

AI 中文总结

本文提出可微混合动作神经控制器(HANC),联合学习切换与连续设定点,在意大利真实区域供热网络模拟中降低30%运行成本,且噪声训练减少硬切换一个数量级。

AI 中文摘要

许多信息物理系统需要结合连续设定点与离散操作决策(如设备切换、模式选择或资源调度)的控制策略。离散动作不可微,阻碍了基于梯度的策略训练,而传统的混合整数规划在线求解成本高昂。本文提出一种混合动作神经控制器(HANC),其中连续分支、分类分支和可微组装层共同生成满足复杂执行器约束的命令。分类决策采用直通Gumbel估计器处理,使策略能够通过时间反向传播在完整闭环轨迹上进行训练。所提出的框架部署在一个具有多个热源和分层热能存储的区域供热网络(DHN)上。其性能在位于意大利RSE SpA的真实DHN模拟中进行了评估。所得策略联合学习切换决策和连续运行设定点。在动态电价下,学习控制器相比基于规则的工业基线降低了30%的运行成本。我们还表明,与确定性直通松弛相比,训练期间注入噪声实现了相似的成本,同时将硬切换减少了一个数量级,并将这一差异归因于所得策略更宽的决策边界。

英文摘要

Many cyber-physical systems require control policies that combine continuous setpoints with discrete operational de- cisions, such as equipment switching, mode selection, or resource scheduling. Discrete actions are not differentiable, which ob- structs gradient-based policy training, while conventional mixed- integer formulations remain costly to solve online. This paper proposes a hybrid-action neural controller (HANC) in which a continuous branch, a categorical branch and a differentiable assembly layer jointly generate commands that satisfy complex actuator constraints by construction. Categorical decisions are handled using a straight-through Gumbel estimator, enabling the policy to be trained by backpropagation through time over full closed-loop rollouts. The proposed framework is deployed on a district heating network (DHN) featuring multiple heat generation units and stratified thermal energy storage. Its performance is evaluated on a simulation of a real DHN located at RSE SpA in Italy. The resulting policy jointly learns switching decisions and continu- ous operating setpoints. Under dynamic electricity pricing, the learned controller reduces operating cost by 30% compared to a rule-based industrial baseline. We also show that, compared with a deterministic straight-through relaxation, injecting noise during training achieves similar cost while reducing hard switching by an order of magnitude, and attribute this difference to the wider decision margins of the resulting policy.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑