arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.11807stat.MLcs.LG

具有多步转移前瞻的近似最优强化学习

Near-Optimal Reinforcement Learning with Multi-Step Transition Lookahead

Corentin Pla, Hugo Richard, Marc Abeille, Vianney Perchet

首次发表
浏览论文内容

中文总结 AI 辅助

本文研究多步转移前瞻的强化学习,证明对任意固定折扣因子精确规划仍NP难,并提出随机多项式时间近似方案,实现与经典方法匹配的遗憾界。

中文摘要 AI 辅助

我们研究了具有转移前瞻(transition look-ahead)的强化学习(RL),其中智能体在决定其行动方案之前,可以观察在采取任何长度为 $\ell$ 的动作序列后将会访问哪些状态。尽管前瞻可以显著提高可实现的性能,但已知具有多步转移前瞻的最优规划是NP难的,然而这一困难性是通过使用任意接近1的折扣因子建立的。因此,尚不清楚该问题对于任何折扣因子是否仍然困难,以及是否仍然可以高效地进行近似最优规划。我们解决了这两个问题。首先,我们证明对于每个固定的有理折扣因子($\gamma\in(0,1)$),精确规划仍然是NP难的。其次,我们为每个固定的前瞻深度引入了一种随机多项式时间近似方案。然后,我们使用乐观主义和方差自适应置信界将我们的方法扩展到未知转移和随机奖励。所得到的算法实现了累积遗憾,其主导项在高达对数因子的情况下与经典表格折扣RL相匹配。因此,尽管具有转移前瞻的精确规划是NP难的,但高效的近似最优规划和学习仍然是可能的。

英文摘要

We study reinforcement learning (RL) with transition look-ahead, where the agent may observe which states would be visited upon playing any sequence of actions before deciding its course of action. Although look-ahead can substantially improve achievable performance, [1] showed that optimal planning with multi-step transition look-ahead is NP-hard. However, this hardness was established using a discount factor close to one. It was therefore unknown whether the problem remains hard for every discount factor, and whether near-optimal planning can nevertheless be performed efficiently. We resolve both questions. First, we show that for every fixed discount factor, exact planning remains NP-hard. Second, we introduce a randomized polynomial-time approximation scheme for every fixed look-ahead depth. Third, we extend our approach to account for unknown transitions. We empirically validate the soundness of our results on the wind-farm storage-control benchmark of [2], showing that our approach, optimally accounting for -step look-ahead information, offers substantially better performance than existing algorithms. [1] Corentin Pla, Hugo Richard, Marc Abeille, Nadav Merlis, Vianney Perchet : On the Hardness of Reinforcement Learning with Transition Look-Ahead [2] Chenbei Lu, Zaiwei Chen, Tongxin Li, Chenye Wu, Adam Wierman : Reinforcement Learning with Imperfect Transition Predictions: A Bellman-Jensen Approach

发表机构

  • CREST(经济与统计研究中心)
  • ENSAE(法国国家统计与经济管理学院)
  • Criteo AI Lab(Criteo人工智能实验室)

机构由 AI 辅助整理,请以论文原文为准。

↑