arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.16888cs.LG

基于Q的变分逆强化学习

Q-based Variational Inverse Reinforcement Learning

Ondrej Bajgar, Peter Tisnikar, Alessandro Abate, Konstantinos Gatsis, Maike Osborne

首次发表
浏览论文内容

中文总结 AI 辅助

提出新型贝叶斯逆强化学习方法QVIRL,兼具可扩展性与不确定性量化能力,在多类任务的学徒学习中表现优异,是首个可从原始像素观测训练的贝叶斯IRL方法。

中文摘要 AI 辅助

安全且有益的AI发展要求系统能按照人类偏好学习和行动,但手动明确指定这些偏好通常不可行。逆强化学习(IRL)通过从专家行为推断作为奖励函数的偏好来应对这一挑战。我们提出Q基变分逆强化学习(QVIRL),这是一种新型贝叶斯IRL方法,主要通过学习最优Q值上的变分分布,从专家演示中恢复奖励的后验分布。与以往方法不同,QVIRL兼具可扩展性与不确定性量化能力,这对安全关键型应用及主动学习十分重要。我们在学徒学习的各类任务中验证了QVIRL的优异性能,包括网格世界、月球着陆器、高速公路环境及两款雅达利游戏,涵盖静态专家数据与主动学习场景。它是首个能从原始像素观测中进行训练的贝叶斯IRL方法。

英文摘要

The development of safe and beneficial AI requires that systems can learn and act in accordance with human preferences. However, explicitly specifying these preferences by hand is often infeasible. Inverse reinforcement learning (IRL) addresses this challenge by inferring preferences, represented as reward functions, from expert behaviour. We introduce Q-based Variational IRL (QVIRL), a novel Bayesian IRL method that recovers a posterior distribution over rewards from expert demonstrations via primarily learning a variational distribution over optimal Q-values. Unlike previous approaches, QVIRL combines scalability with uncertainty quantification, important for safety-critical applications as well as active learning. We demonstrate QVIRL's strong performance in apprenticeship learning across various tasks, including gridworlds, Lunar Lander, the Highway Environment, and two ATARI games both with static expert data and with active learning. It is the first method for Bayesian IRL that demonstrates training from raw pixel observations.

↑