发表机构
Shanghai Artificial Intelligence Laboratory; University of Auckland(上海人工智能实验室; 奥克兰大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出无循环逆强化学习(LFIRL),通过扩散策略编码的动作梯度结构,将奖励学习转化为顺序价值恢复问题,消除策略优化循环,实现2-3倍加速并保持高质量奖励恢复。
AI 中文摘要
逆强化学习(IRL)旨在恢复能够解释专家示范的奖励函数。现有的IRL方法通常依赖于双层优化过程,在奖励学习和策略优化之间交替进行,导致巨大的计算负担和训练不稳定性。在这项工作中,我们引入了一种不同的路径,通过利用扩散策略完全消除策略优化。我们的关键见解是,扩散策略编码了最优软Q函数的动作梯度结构,使得奖励学习可以被视为一系列价值恢复问题,从而允许我们绕过先前IRL方法中固有的奖励-策略循环。具体来说,我们的方法分三个阶段进行:(I)通过动作梯度匹配恢复最优软Q函数,并以受Gumbel回归启发的方式估计相应的软价值函数(Q值的LogSumExp);(II)通过推断状态相关的偏移来校准这些软价值;(III)通过强制执行Bellman一致性来提取奖励。这导致了无循环逆强化学习(LFIRL),一种完全离线的算法,以简单、无循环和顺序的方式运行。LFIRL易于实现,显著提高了训练效率,同时保持了强大的奖励恢复性能。在实验上,在Maze、Franka Kitchen、Adroit Hand Pen和Push-T基准测试中,LFIRL实现了比最快基线2-3倍的加速,同时在奖励恢复质量上匹配或超越了最先进的方法。
英文摘要
Inverse Reinforcement Learning (IRL) aims to recover a reward function that explains expert demonstrations. Existing IRL methods typically rely on a bi-level optimization procedure that alternates between reward learning and policy optimization, leading to substantial computational burden and training instability. In this work, we introduce a different route that eliminates policy optimization entirely by leveraging diffusion policies. Our key insight is that a diffusion policy encodes the action-gradient structure of the optimal soft Q function, enabling reward learning to be cast as a sequence of value recovery problems, thereby allowing us to bypass reward-policy loops inherent in prior IRL methods. Specifically, our method proceeds in three stages: (I) recovering the optimal soft Q function via action-gradient matching and estimating the corresponding soft value function (LogSumExp of Q values) in a way inspired by Gumbel regression; (II) calibrating these soft values by inferring a state-dependent offset; (III) extracting the reward by enforcing Bellman consistency. This leads to Loop-Free Inverse Reinforcement Learning (LFIRL), a fully offline algorithm that operates in a simple, loop-free, and sequential manner. LFIRL is simple to implement and significantly improves training efficiency while maintaining strong reward recovery performance. Empirically, across Maze, Franka Kitchen, Adroit Hand Pen, and Push-T benchmarks, LFIRL achieves 2-3x speedup over the fastest baselines, while matching or surpassing state-of-the-art methods in reward recovery quality.
CommentsAccepted to NeurIPS 2026. 20 pages, 7 figures