arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.17031eess.SYcs.SY

基于具有量子采样特征的混合软演员-评论家的可行性感知安全约束机组组合

Feasibility-Aware Security-Constrained Unit Commitment via Hybrid Soft Actor-Critic with Quantum-Sampled Features

George Dimas, Amin Masoumi, Mert Korkali

首次发表
浏览论文内容

中文总结 AI 辅助

研究针对安全约束机组组合计算成本高的问题,提出三层混合框架,用伯努利混合软演员-评论家策略、量子采样辅助通道及原生SCUC混合整数线性规划求解,评估多节点案例,发现有用承诺信息的量是影响可扩展性的主要限制。

中文摘要 AI 辅助

安全约束机组组合(SCUC)在多周期范围内耦合二进制承诺、经济调度、备用和网络安全,对于实际系统规模而言,精确求解计算成本高昂。本文提出了一个三层混合框架,其中伯努利混合软演员-评论家(HSAC)策略提出每小时的承诺,量子采样辅助通道增强状态,并且在仅执行有限子集的承诺二进制数之后,原生SCUC混合整数线性规划恢复调度和安全变量。该方法因此与求解器兼容,而非对精确优化的端到端替代。我们形式化了从SCUC到强化学习的接口,推导了由固定上限引起的时间覆盖范围,并评估了14、57和118节点的基准案例。结果表明,在14节点案例中恢复稳定且成本低,最佳恢复调度达到全时段最优;在57节点案例中筛选拒绝率非常低;而在118节点案例中,一旦执行上限不再跨越完整的承诺期,就会出现明显的覆盖瓶颈。因此,该研究确定了在探索性伯努利演员和小执行上限下,到达恢复模型的有用承诺信息的量是控制可扩展性的主要限制。

英文摘要

Security-constrained unit commitment (SCUC) couples binary commitment, economic dispatch, reserves, and network security over a multiperiod horizon, making an exact solution computationally expensive for realistic system sizes. This paper proposes a three-layer hybrid framework in which a Bernoulli hybrid soft actor-critic (HSAC) policy proposes hourly commitments, a quantum-sampled auxiliary channel augments the state, and a native SCUC mixed-integer linear program recovers dispatch and security variables after only a limited subset of commitment binaries is enforced. The method is therefore solver-compatible rather than an end-to-end replacement for exact optimization. We formalize the SCUC-to-reinforcement-learning interface, derive the temporal coverage induced by the fixed cap, and evaluate the 14- 57- and 118-bus benchmark cases. The results show stable, low-cost recovery in the 14-bus case, where the best recovered schedule attains the full-horizon optimum; a very low screen-rejection rate in the 57-bus case; and a clear coverage bottleneck in the 118-bus case once the enforcement cap no longer spans a complete commitment period. The study, therefore, identifies the amount of useful commitment information that reaches the recovery model, under an exploratory Bernoulli actor and a small enforcement cap, as the dominant limitation that governs scalability

↑