发表机构
The University of Tokyo; The University of Osaka; RIKEN(东京大学; 大阪大学; 理化学研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对带对抗损失和多臂老虎机反馈的在线分段表格型MDP,用正则化Q函数缩小策略优化的视界差距,改进了遗憾界的视界依赖关系,还扩展到对抗线性混合MDP。
AI 中文摘要
我们研究带有对抗损失和多臂老虎机反馈的在线分段表格型马尔可夫决策过程(MDP)的策略优化问题。策略优化在每个状态处局部更新策略,避免在占用测度多面体上进行优化,但现有策略优化的遗憾界比基于占用测度的算法大一个视界H的因子。我们通过使用正则化Q函数缩小这一差距,该函数可联合控制所有状态-动作对的局部更新稳定性,而非在每个状态单独控制。所得算法在转移已知时达到高概率遗憾界\\(\widetilde O(\sqrt{HS(H+A)T})\\),在转移未知时达到\\(\widetilde O(HS\sqrt{AT})\\),其中S为状态数,A为动作数,T为片段数。两个界均改进了现有策略优化界的视界依赖关系,后者匹配最知名的界。我们进一步将算法扩展到对抗线性混合MDP,获得相同的视界依赖改进。
英文摘要
We consider policy optimization for online episodic tabular Markov decision processes (MDPs) with adversarial losses and bandit feedback. Policy optimization updates the policy locally at each state and avoids optimization over the occupancy-measure polytope, but its existing regret bounds are larger by a factor of the horizon $H$ than those of occupancy-measure-based algorithms. We close this gap by using regularized $Q$-functions, which allow us to control the stability of the local updates jointly over all state-action pairs rather than separately at each state. The resulting algorithm attains high-probability regret bounds of $\widetilde O(\sqrt{HS(H+A)T})$ for known transitions and $\widetilde O(HS\sqrt{AT})$ for unknown transitions, where $S$ is the number of states, $A$ the number of actions, and $T$ the number of episodes. Both bounds improve the horizon dependence of existing policy optimization bounds, and the latter matches the best-known bound. We further extend the algorithm to adversarial linear-mixture MDPs and obtain the same improvement in the horizon dependence.
Comments17 pages, 2 tables