发表机构
Shanghai Jiao Tong University; SIMIS; Stanford University(上海交通大学; 上海数学与交叉学科研究院; 斯坦福大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究证明霍华德策略迭代在确定性折扣MDP上具有指数迭代下界,排除其强多项式性,并与单纯形法形成指数分离,揭示算法无政府状态的价格。
AI 中文摘要
我们为霍华德策略迭代在确定性折扣马尔可夫决策过程上建立了关于状态数量的指数迭代下界,其中每个状态最多有两个动作。这排除了当折扣因子作为输入一部分时霍华德策略迭代的强多项式性,并产生了与使用丹齐格枢轴规则的单纯形法的指数分离,后者被证明在该类问题上具有强多项式性。即使每个奖励被限制为对数位长度,我们也获得了拉伸指数迭代下界。霍华德的去中心化与同时自私改进和丹齐格在所有状态中协调选择具有最大增益的单一动作之间的差距,揭示了算法无政府状态的“价格”。
英文摘要
We establish an exponential iteration lower bound in the number of states for Howard's policy iteration on deterministic discounted Markov decision processes, with at most two actions per state. This rules out strong polynomiality of Howard's policy iteration when the discount factor is part of the input and yields an exponential separation from the simplex method with Dantzig's pivoting rule, which is proved to be strongly polynomial on this class. Even when each reward is restricted to logarithmic bit length, we obtain a stretched-exponential iteration lower bound. The gap between Howard's decentralized and simultaneous selfish improvements and Dantzig's coordinated selection of a single action with the largest gain across all states reveals a ``price'' of algorithmic anarchy.