arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.22546eess.SYcs.SY

从电池能量管理系统中的异构多时域时间序列学习控制策略

Learning Control Policies from Heterogeneous Multi-Horizon Time Series in Battery Energy Management Systems

Sheng Yin, Vivek Teja Tanjavooru, Holger Hesse, Christoph Goebel

首次发表
浏览论文内容

中文总结 AI 辅助

R2D是一种端到端模仿学习框架,直接从异构多时域时间序列学习电池控制策略,通过LSTM编码器和MILP专家行为克隆,在五个工业站点上达到全局最优的62%-77%,优于MPC和RL基准,且电池退化接近专家水平。

中文摘要 AI 辅助

本文介绍了表示到决策(R2D),一种端到端的模仿学习框架,通过模块化时间特征提取器(TFEs)和共享潜在表示,将异构多时域时间序列输入直接映射到电池控制决策,无需显式的负荷和光伏预测步骤。虽然准确的预测能提高预测质量,但在预测-然后-优化的流程中,最优控制性能仍无法得到保证;标准的强化学习(RL)由于短窗口观测而缺乏长时间范围的时间感知,即使将预测信号作为额外输入也是如此。R2D提供了不同的视角:不是先预测后决策,而是直接从原始时间输入学习决策,通过行为克隆(BC)以老化感知的混合整数线性规划(MILP)专家为锚定控制最优性。在五个工业站点的高保真电热电池模拟中,与六个控制器进行基准测试,R2D达到了全局先知最优的62%-77%,优于测试的模型预测控制(MPC)和RL基准,并且在所有五个站点上产生的电池退化几乎与MILP教师相同。对时间骨干网络、模型大小、专家公式、老化成本权重和时域配置的全面消融研究证实,具有15分钟单步控制时域的LSTM编码器提供了最稳健且可部署的配置,跨站点和单因素分布外测试表明,对未见配置的泛化是依赖于配置的,在高活动站点上训练的策略迁移最可靠。

英文摘要

This paper introduces Representation-to-Decision (R2D), an end-to-end imitation learning framework that maps heterogeneous multi-horizon time-series inputs directly to battery control decisions through modular Temporal Feature Extractors (TFEs) and a shared latent representation, without an explicit load and PV forecasting step. While accurate forecasting improves prediction quality, optimal control performance remains unguaranteed in prediction-then-optimization pipelines; standard Reinforcement Learning (RL) lacks the long-horizon temporal awareness due to short-window observations, even when forecast signals are available as additional inputs. R2D offers a different perspective: rather than forecasting first and deciding second, it learns to decide directly from raw temporal inputs, with control optimality anchored by an aging-aware Mixed-Integer Linear Programming (MILP) expert through Behavior Cloning (BC). Benchmarked against six controllers on a high-fidelity electro-thermal battery simulation across five industrial sites, R2D achieves 62--77\% of the global clairvoyant optimum, outperforms the tested Model Predictive Control (MPC) and RL benchmarks, and yields battery degradation nearly identical to its MILP teacher across all five sites. Comprehensive ablation studies over temporal backbone, model size, expert formulation, aging-cost weighting, and horizon configuration confirm that an LSTM encoder with a 15-minute single-step control horizon provides the most robust and deployment-ready configuration, and cross-site and single-factor out-of-distribution tests show that generalization to unseen profiles is profile-dependent, with policies trained on high-activity sites transferring most reliably.

发表机构

  • Technical University of Munich(慕尼黑工业大学)
  • Kempten University of Applied Sciences(凯姆滕应用技术大学)

机构由 AI 辅助整理,请以论文原文为准。

↑