arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.11052cs.LGstat.ML

用于逆强化学习的高效超梯度下降

Efficient Hypergradient Descent for Inverse Reinforcement Learning

Nikita Sevriukov, Anna Barabanova, Uliana Gagarina, Karina Ivanova, Sofiia Kasaeva, Ilya Levin, Marina Sheshukova

首次发表
浏览论文内容

中文总结 AI 辅助

该研究针对逆强化学习双层优化的计算挑战,利用策略费舍尔信息矩阵的特性设计结构化超梯度,通过流式频谱草图近似逆费舍尔向量乘积,在控制环境中实现了高效且性能良好的IRL方法。

中文摘要 AI 辅助

逆强化学习(IRL)旨在恢复一个奖励函数,使得基于该奖励函数得到的策略能够复现专家演示中观察到的行为。一种自然的方法是将IRL表述为双层优化问题,其中内层对应于在学习到的奖励下的策略优化,外层则衡量诱导策略与专家数据之间的差异。然而,这种表述在实际应用中计算难度较大,因为外层更新需要涉及内层目标的逆海森向量乘积的超梯度。我们通过证明在内层最优解处,内层目标的海森与策略的费舍尔信息矩阵成比例,从而得到了一种与自然超梯度下降密切相关的结构化基于费舍尔的超梯度,以此解决这一挑战。为解决与大型费舍尔矩阵相关的可扩展性瓶颈,我们使用流式频谱草图近似所需的逆费舍尔向量乘积,避免显式构建费舍尔矩阵。我们在离散和连续控制环境中,将我们的方法与一阶随机双层基线进行了评估。结果表明,该方法具有有竞争力的策略性能和出色的奖励排序质量,同时费舍尔草图降低了曲率存储复杂度,并且相较于显式费舍尔求解器可提高计算效率。

英文摘要

Inverse reinforcement learning (IRL) aims to recover a reward function under which the resulting policy reproduces the behavior observed in expert demonstrations. A natural approach is to formulate IRL as a bilevel optimization problem, in which the inner level corresponds to policy optimization under the learned reward and the outer level measures the discrepancy between the induced policy and expert data. However, this formulation is computationally challenging in practice because the outer update requires a hypergradient involving an inverse-Hessian-vector product for the inner objective. We address this challenge by showing that, at the inner optimum, the Hessian of the inner objective is proportional to the Fisher information matrix of the policy, yielding a structured Fisher-based hypergradient closely related to Natural Hypergradient Descent. To address the resulting scalability bottleneck associated with large Fisher matrices, we approximate the required inverse-Fisher-vector product using a streaming spectral sketch, avoiding explicit construction of the Fisher matrix. We evaluate our approach against a first-order stochastic bilevel baseline across discrete- and continuous-control environments. The results demonstrate competitive policy performance and strong reward-ranking quality, while Fisher sketching reduces curvature-storage complexity and can improve computational efficiency relative to an explicit Fisher solver.

发表机构

  • HSE University(高等经济大学)

机构由 AI 辅助整理,请以论文原文为准。

↑