arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

表格型马尔可夫决策过程中具有转移过渡动态的混合强化学习统一算法框架

A Unified Algorithmic Framework for Hybrid Reinforcement Learning in Tabular MDPs with Shifted Transition Dynamics

Zheshun Wu, Renjie Zheng, Jinhang Zuo, Zenglin Xu, Fang Kong

arXiv 2607.25207首次发表:更新:

发表机构

Southern University of Science and Technology; Harbin Institute of Technology, Shenzhen; City University of Hong Kong; Fudan University; Shanghai Academy of AI for Science(南方科技大学; 哈尔滨工业大学(深圳); 香港城市大学; 复旦大学; 上海人工智能研究院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究表格型MDP中混合强化学习,提出含MIN-UCB-VI和MAX-LCB-VI的统一算法框架,利用细粒度偏差信息有效利用离线数据,给出理论保证并通过实验验证,解决离线数据因转移动态而无效的问题。

AI 中文摘要

本文研究表格型马尔可夫决策过程(MDP)中的混合强化学习设置,智能体旨在通过结合与目标环境的在线交互和来自源环境的离线数据来学习最优策略。关键挑战是离线数据可能来自具有转移过渡动态的过时环境,使历史数据的简单整合无效。为此,我们提出统一算法框架,包含MIN-UCB-VI用于最小化遗憾值和MAX-LCB-VI用于识别最佳策略。两种算法利用细粒度偏差信息在一般转移下更有效利用离线数据。我们给出理论保证,包括遗憾值和次优差距的实例相关和独立上界。还建立匹配下界证明方法最优性,并通过大量实验验证理论结果。

英文摘要

This paper investigates a hybrid reinforcement learning setting in tabular Markov Decision Processes (MDPs), where an agent aims to learn an optimal policy by combining online interactions with a target environment and offline data from a source environment. A central challenge is that offline data may be collected from outdated environments with shifted transition dynamics, making naive integration of historical data ineffective. To address this, we propose a unified algorithmic framework featuring two algorithms: MIN-UCB-VI for regret minimization and MAX-LCB-VI for best policy identification. Both algorithms leverage fine-grained bias information to more effectively exploit offline data under general transition shifts. We provide theoretical guarantees for our framework, including both instance-dependent and independent upper bounds on regret and sub-optimality gap. Furthermore, we establish matching lower bounds to demonstrate the optimality of our approach and validate our theoretical findings through extensive experiments.

Comments59 pages, 3 figures, and 2 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑