arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

面向目标条件强化学习的世界模型策略仲裁器

World-Model Policy Arbiter for Goal-Conditioned Reinforcement Learning

Junwei Quan, Evgenii Opryshko, Nicholas Rhinehart, Igor Gilitschenski

arXiv 2610.10932首次发表:更新:

发表机构

University of Toronto; Vector Institute(多伦多大学; 矢量研究所)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对目标条件强化学习中单一策略表现不佳的问题,提出WMPA框架,通过世界模型评估冻结策略并选最优执行,在18个数据集上提升了宏平均成功率。

AI 中文摘要

离线目标条件强化学习(GCRL)已产生多种目标达成算法,但没有一种算法能在所有环境、所有目标甚至同一任务的不同阶段都表现最佳。我们不仅部署表现最好的策略,而是探究能否将一组冻结的目标条件策略作为组合策略协同使用,在每个状态下决定由哪个策略执行动作。在每个状态下选择策略并非易事:各策略自身的值函数无法直接比较,它们可能使用不同的尺度,且部分策略没有值函数。我们需要根据各策略可能到达的状态来评判它们,即便一次只能执行一个策略;同时还需避免切换过于频繁导致控制不稳定。为应对这些挑战,我们提出世界模型策略仲裁器(WMPA),这是一种测试时框架,给定一组冻结策略作为输入,在学习到的状态空间世界模型中展开每个冻结策略,用共享的目标条件值函数评估想象的未来,在下次仲裁(策略选择)前,让得分最高的策略执行一段短承诺间隔。WMPA假设可获取一组冻结的目标条件策略,既不需要重新训练策略,也不需要特定任务的特权知识。在涵盖迷宫导航、立方体、场景及谜题操作的18个基于状态的数据集上,采用官方OGBench评估协议,WMPA将宏平均成功率从按每个数据集选最优策略的44%提升至58%,在12个数据集上取得统计显著提升,其中cube-double-play提升33个百分点,scene-play提升36个百分点。

英文摘要

Offline goal-conditioned reinforcement learning (GCRL) has produced a diverse set of goal-reaching algorithms, yet no single algorithm performs best across environments, goals, and even different phases of the same task. Rather than deploying only the best-performing policy, we ask whether a set of frozen goal-conditioned policies can be used collectively as a portfolio, deciding at every state which policy should act. Choosing a policy at each state is not straightforward. The policies' own value functions cannot be compared directly: they may use different scales, and some policies have no value function. We need to judge each policy by the states it is likely to reach, even though we can execute only one policy at a time. We also need to avoid switching so often that control becomes unstable. To address these challenges, we introduce World-Model Policy Arbiter (WMPA), a test-time framework that, given a set of frozen policies as input, rolls out each frozen policy in a learned state-space world model, evaluates the imagined futures with a shared goal-conditioned value function, and executes the highest-scoring policy for a short commitment interval before the next round of arbitration (policy selection). WMPA assumes access to a bank of frozen goal-conditioned policies and requires neither policy retraining nor privileged task-specific knowledge. Under the official OGBench evaluation protocol on 18 state-based datasets spanning maze navigation as well as cube, scene, and puzzle manipulation, WMPA improves the macro-average success rate from the 44% achieved by the best policy selected per dataset to 58%, with statistically significant gains on 12 datasets. These gains include +33 percentage points on cube-double-play and +36 percentage points on scene-play.

Comments22 pages, 3 figures, 14 tables. Project page: https://junwei0102.github.io/wmpa-webpage/ Code: https://github.com/junwei0102/World-Model-Policy-Arbiter

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑