RL-PaO:不确定性决策中的预测即行动
RL-PaO: Prediction as Action in Decision Making under Uncertainty
浏览论文内容
中文总结 AI 辅助
RL-PaO提出将预测视为行动的强化学习框架,统一系统建模、优化与决策执行,在日前能源调度中实现平均10%的成本降低,并提供了策略演化与成本-准确性权衡的可解释性分析。
中文摘要 AI 辅助
不确定性下的决策通常依赖于预测参数,然而准确的预测并不一定能带来良好的运营决策。将预测与下游优化对齐,需要从这些预测所引发的决策后果中学习。我们提出了RL-PaO,一个将系统建模、优化和决策执行整合到单一环境中的强化学习框架。这产生了一个马尔可夫决策过程,其中预测被视为行动:它改变环境以产生后续的上下文和奖励,从而明确地将预测误差与实际成本对齐,并且学习最优策略不需要对黑盒求解器进行微分。我们使用真实历史数据在日前能源调度上评估了RL-PaO。在测试年份中,RL-PaO在非神谕基线中实现了最低的年度成本,平均降低了10%的成本。此外,RL-PaO能够进行进一步分析,从策略演化和成本-准确性权衡的角度提供强大的可解释性。
英文摘要
Decision-making under uncertainty often relies on predicted parameters, yet accurate prediction does not necessarily lead to good operational decisions. Aligning prediction with downstream optimization requires learning from the consequences of the decisions those predictions induce. We introduce RL-PaO, a reinforcement learning framework that integrates system formulation, optimization, and decision execution into a single environment. This yields a Markov decision process in which prediction is regarded as action: it shifts the environment to produce subsequent context and reward that explicitly aligns prediction error with realized cost, and learning the optimal policy does not require differentiating through the black-box solver. We evaluate RL-PaO on day-ahead energy scheduling using real historical data. On the test year, RL-PaO achieves the lowest annual cost among the non-oracle baselines, achieving on average $10\%$ cost reduction. Moreover, RL-PaO is capable of further analyses to provide strong interpretability both from the policy evolution perspective and the cost-accuracy trade-off.
发表机构
- Great Bay University(大湾区大学)
- The University of Osaka(大阪大学)
- Aalto University(阿尔托大学)
机构由 AI 辅助整理,请以论文原文为准。