发表机构
IDLab Ghent university – imec; Imperial College London(IDLab 根特大学 – imec; 帝国理工学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对多阶段随机MPC的场景树构建,提出基于强化学习的面向控制的方法,在电池套利问题上验证其收益、鲁棒性及尾部风险特性均优于传统方法。
AI 中文摘要
多阶段随机模型预测控制(MPC)通过在场景树(一种由采样预测值构建的未来结果的有限分支近似)上进行优化来处理不确定性。传统构建场景树的方法聚焦于匹配潜在概率分布,例如基于Wasserstein的场景约简,但分布精度提升未必带来更优的控制性能。本文提出一种面向控制的方法,直接从场景树对下游决策的影响中学习其构建方式。在固定树拓扑的前提下,我们将场景树构建建模为采样场景到叶节点的序贯分配问题,该分配由场景集上的基于注意力的策略参数化,并以闭环控制收益为目标通过强化学习进行训练;训练过程通过利用已实现未来轨迹的不对称评论家实现稳定。我们在风险规避型电池套利问题上评估该方法,在一系列预测集规模下,所学习的构建方式始终实现最高收益,优于经典的前向、后向约简方法及确定性等价(单轨迹预测)控制。所学习策略在挑战性案例上表现出更强鲁棒性,始终展现更优的尾部风险特性。对生成树的分析表明,本文方法构建出紧凑、选择性分支的结构,既能捕捉高影响事件,又能保持多数轨迹近乎确定性。这些发现强调场景树的价值关键取决于其所支持的决策,并提供了一个仅基于闭环控制优化信号训练场景树构建器的有效框架。
英文摘要
Multistage stochastic model predictive control (MPC) handles uncertainty by optimizing over a scenario tree, a finite branching approximation of future outcomes constructed from sampled forecasts. To build such a tree, conventional methods focus on matching the underlying probability distribution---e.g., via Wasserstein-based scenario reduction---but improved distributional accuracy does not necessarily yield better control performance. We propose a control-oriented approach that learns scenario tree construction directly from its impact on downstream decisions. Fixing the tree topology, we formulate tree construction as a sequential assignment of sampled scenarios to leaves. This assignment is parameterized by an attention-based policy over the scenario set and trained using reinforcement learning, with closed-loop control profit as the objective. Training is stabilized by an asymmetric critic that leverages realized future trajectories. We evaluate the method on a risk-averse battery arbitrage problem. Across a range of forecast set sizes, the learned construction consistently achieves the highest profit, outperforming classical forward and backward reduction methods and certainty-equivalent (single-trajectory forecast) control. The learned policy also exhibits greater robustness on challenging instances, consistently demonstrating better tail-risk characteristics. Analysis of the resulting trees indicates that our method constructs compact, selectively branching structures that capture high-impact events while keeping most trajectories nearly deterministic. These findings highlight that the value of a scenario tree depends critically on the decisions it supports, and provide an effective framework to train scenario tree constructors merely based on the closed-loop control optimization signal.