arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.17966math.OCcs.SYeess.SYmath.PR

BAR-SOT方法:长期平均成本控制作为随机最优自输运

The BAR-SOT Method: Long-term Average Cost Control as Stochastic Optimal Self-Transport

Sharan Srinivasan, Berke M. Turkay, Harsha Honnappa

首次发表
浏览论文内容

中文总结 AI 辅助

提出BAR-SOT方法,将长期平均成本控制重构为有限时域随机最优自输运问题,通过三种等价表述及神经对偶参数化,在多个基准上超越强化学习基线和优先级启发式。

中文摘要 AI 辅助

我们将马尔可夫跳跃过程的平均成本(遍历)控制重新表述为一个有限时域随机最优输运(SOT)问题,该问题联合优化受控演化及其起始并返回的边缘分布(即自输运)。将边缘流与受控生成器联系起来的约束是基本伴随关系(BAR),因此我们将所得问题称为BAR-SOT。对于任意时间范围T>0,其最优值是长期平均成本率的常数缩放。我们给出了三种等价表述(通过受控过程、Fokker-Planck约束和松弛边缘测度),并表明最优对偶是一个平稳势加上时间线性项,其斜率等于该比率。对控制施加相对熵惩罚产生了一个成本倾斜的Schrödinger桥问题,当每个转移率都被控制时,可通过Sinkhorn型迭代计算,其值在惩罚消失时收敛到未正则化的最优值。我们为有限马尔可夫决策过程发展该理论,然后推广到一般受控马尔可夫跳跃过程。对偶的神经参数化,作为物理信息神经网络(PINN)训练,编码了Harrison的等效工作负载表述(见Harrison 2000),与Dai和Gluzman(2022)的强强化学习基线相匹配;对次主导横向修正的拟合进行数值条件化后,改进了该基线和最佳优先级启发式。在输入排队交换机上,对偶的图注意力参数化改进我们所知的最强匹配启发式,且参数数量不随交换机规模增长。

英文摘要

We reformulate average-cost (ergodic) control of a Markov jump process as a finite-horizon stochastic optimal transport (SOT) problem that jointly optimizes over the controlled evolution and the marginal law from which it starts and returns to (i.e., a self-transport). The constraint relating the marginal flow to the controlled generator is the basic adjoint relationship (BAR), so we call the resulting problem BAR-SOT. For any time horizon T>0, its optimal value is a constant scaling of the long-run average-cost rate. We give three equivalent formulations (through controlled processes, a Fokker--Planck constraint, and relaxed marginal measures) and show that the optimal dual is a stationary potential plus a term linear in time, with slope equal to the rate. A relative-entropy penalty on the control yields a cost-tilted Schrodinger bridge problem, computable by a Sinkhorn-type iteration when every transition rate is controlled, whose value converges to the unregularized optimum as the penalty vanishes. We develop the theory for finite Markov decision processes and then for general controlled Markov jump processes. A neural parametrization of the dual, trained as a physics-informed neural network (PINN), that encodes Harrison's equivalent-workload formulation (see Harrison 2000) matches the strong reinforcement-learning baseline of Dai and Gluzman (2022); numerically conditioning the fit of a sub-dominant transverse correction then improves on both that baseline and the best priority heuristic. On the input-queued switch, a graph-attention parametrization of the dual improves on the strongest matching heuristic we are aware of, with a parameter count that does not grow with the switch size.

发表机构

  • Purdue University(普渡大学)

机构由 AI 辅助整理,请以论文原文为准。

↑