发表机构
Universitat Pompeu Fabra(庞培法布拉大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出鲁棒后继特征,在线性MDP假设下统一了跨奖励函数与转移核的泛化,推导了广义策略改进的性能退化界,并在网格基准上验证其优于仅关注单一因素的迁移方法。
AI 中文摘要
强化学习中的泛化指的是智能体在一组任务上训练后,能够在未见过的任务上执行接近最优策略的能力。基于后继表示的开创性工作以及函数逼近的进一步改进,传统上强化学习中的迁移聚焦于泛化到仅在奖励函数上有所不同的任务。在后继表示引入十年后,鲁棒强化学习从运筹学领域的多篇文章中同时涌现。在鲁棒强化学习中,转移核是未知的,目标是在这种不确定性下最大化期望奖励。我们的工作通过鲁棒后继特征统一了这两种范式,在线性马尔可夫决策过程的假设下,鲁棒后继特征能够同时跨奖励函数和转移核进行泛化。我们推导了广义策略改进的一个界,该界明确量化了性能如何随转移核之间的不匹配而下降,并在动力学共享时恢复了现有的后继特征保证。最后,鲁棒后继特征的泛化能力在多个基于网格的基准上得到验证,并与先前仅关注奖励或仅关注转移核的替代方法进行了比较。
英文摘要
Generalization in Reinforcement Learning (RL) refers to the ability to execute close-to-optimal policies in unseen tasks after the agent has been trained on a different set of tasks. Building on the seminal work of the successor representation and further adaptations with function approximation, Transfer in RL has traditionally focused on generalizing to tasks that only differ in the reward function. A decade after the introduction of the successor representation, Robust RL emerged simultaneously from several articles in the field of operations research. In Robust RL, the transition kernel is unknown, and the goal is to maximize the expected reward under this uncertainty. Our work unifies these two paradigms through robust successor features, which generalize across both the reward function and the transition kernel, under the assumption that tasks are linear Markov Decision Processes. We derive a bound on Generalized Policy Improvement (GPI) that explicitly quantifies how performance degrades with the mismatch between transition kernels, recovering existing successor-feature guarantees when dynamics are shared. Finally, the generalization capabilities of robust successor features are validated on several grid-based benchmarks and compared to previous alternatives that focus solely on either the reward or the transition kernel.
Comments10 pages, 3 figures, to be published in EWRL 2026