发表机构
Department of Computer Science, University of Pisa, Italy; Department of Economics and Statistics, University of Salerno, Italy; Department of Statistics, Computer Science, Applications ”G.Parenti”, University of Florence, Italy(比萨大学计算机科学系; 萨勒诺大学经济学与统计系; 佛罗伦萨大学统计学、计算机科学与应用G.Parenti系)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出LF-IBIS算法,结合近似贝叶斯计算与迭代批量重要性采样,在似然函数不可得时实现强化学习中的全贝叶斯推断,量化策略不确定性以平衡探索-利用。
AI 中文摘要
强化学习(RL)是一个序贯决策框架,智能体通过与环境的交互学习最优策略,以最大化累积奖励。在RL方法中,贝叶斯强化学习(BRL)通过利用关于环境的先验知识和序贯信念更新,解决了与数据稀缺相关的常见实际挑战。然而,大多数BRL方法需要显式的似然函数,这在现实场景中通常不可访问或难以处理。我们提出了无似然迭代批量重要性采样(LF-IBIS),一种新的BRL算法,它在新交互可用时在线更新智能体的信念。通过将近似贝叶斯计算与迭代批量重要性采样相结合,LF-IBIS能够在环境动态不由显式或可处理的似然函数描述的情况下进行全贝叶斯推断。该方法产生关于环境参数和最优策略的近似后验分布,提供了对策略不确定性的量化,这对于探索-利用权衡的贝叶斯处理非常有用。我们在临床试验的反应自适应随机化模拟研究中测试了该方法,其中闭式后验可用于验证。额外的实验针对后验无闭式形式的情况,并展示了基于最优策略后验分布的在线策略更新。
英文摘要
Reinforcement Learning (RL) is a sequential decision-making framework in which an agent learns optimal policies through interaction with an environment by maximizing cumulative rewards. Among RL methods, Bayesian Reinforcement Learning (BRL) addresses common practical challenges related to data scarcity by leveraging prior knowledge about the environment and sequential belief updates. However, most BRL approaches require an explicit likelihood function, which is frequently inaccessible or intractable in real-world settings. We propose Likelihood-Free Iterated Batch Importance Sampling (LF-IBIS), a novel algorithm for BRL that updates the agent's beliefs online as new interactions become available. By combining Approximate Bayesian Computation with Iterated Batch Importance Sampling, LF-IBIS enables full Bayesian inference in settings where the environment dynamics are not described by an explicit or tractable likelihood. The method yields approximate posterior distributions over both environment parameters and optimal policies, providing a quantification of policy uncertainty useful for a Bayesian treatment of the exploration-exploitation trade-off. We test the method on a simulation study in response-adaptive randomization in clinical trials, where closed-form posteriors enable validation. Additional experiments address settings where the posterior has no closed form and illustrate online policy updating based on the posterior distribution of the optimal policy.
Comments37 pages, 12 figures, 4 tables