arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

微电网能量协调中联邦强化学习的约束感知聚合

Constraint-Aware Aggregation for Federated Reinforcement Learning in Microgrid Energy Coordination

Usman Haider, Karl Mason

arXiv 2607.12763首次发表:更新:

发表机构

School of Computer Science, University of Galway(戈尔韦大学计算机科学学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究微电网能量协调中联邦强化学习的约束感知聚合,提出含本地性能和约束违反的聚合规则,基于惩罚的规则效果最佳,实验表明轻量级聚合策略能在保留通信协议时提高联邦强化学习的经验安全性。

AI 中文摘要

联邦强化学习(FedRL)可在不共享原始本地数据的情况下协调分布式能源资源,但像FedAvg这样的标准聚合方法未考虑系统级约束,常导致不安全的全局行为。本文研究分布式能源协调中联邦强化学习的约束感知聚合。我们提出将本地性能和估计的约束违反纳入服务器端更新的聚合规则。其中,一个简单的基于惩罚的规则$w_i \propto R_i - \alpha V_i$在奖励和安全之间始终提供最可靠的权衡,无需对偶优化或修改本地训练。我们在DairyGridEnv上评估方法,该基准模拟多个农场在随机需求和共享电网容量约束下协调电池存储,并使用芬兰的实际负载驱动需求曲线和德国FIELD数据集进一步评估鲁棒性。在多个种子设置下,基于惩罚的聚合在合成和实际负载驱动设置中相对于FedAvg显著减少违反并提高奖励。组合奖励 - 违反方案通过$\lambda$展现可调权衡,但稳定性较差。这些结果表明轻量级聚合策略可在保留标准通信协议的同时大幅提高联邦强化学习中的经验安全性。

英文摘要

Federated Reinforcement Learning (FedRL) enables coordination of distributed energy resources without sharing raw local data, but standard aggregation methods such as FedAvg do not account for system-level constraints, often leading to unsafe global behavior. In this work, we study constraint-aware aggregation for federated reinforcement learning in distributed energy coordination. We propose aggregation rules that incorporate both local performance and estimated constraint violation into the server-side update. Among these, a simple penalty-based rule, $w_i \propto R_i - αV_i$, consistently provides the most reliable trade-off between reward and safety, without requiring dual optimization or modifications to local training. \textcolor{black}{We evaluate our approach on DairyGridEnv, a benchmark modeling multiple farms coordinating battery storage under stochastic demand and a shared grid capacity constraint, and further assess robustness using real load-driven demand profiles from Finland and the German FIELD dataset. Across multiple seeds, penalty-based aggregation substantially reduces violations while improving reward relative to FedAvg in both synthetic and real load-driven settings.} A combined reward-violation scheme exposes a tunable trade-off via $λ$, but is less stable. These results demonstrate that lightweight aggregation strategies can substantially improve empirical safety in federated reinforcement learning while preserving standard communication protocols.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑