arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

正则化的可证明优势:对抗模仿学习的快速收敛率

Provable Benefits of Regularization: Fast Rates for Adversarial Imitation Learning

Hanbin Zhou, Shangzhe Li, Alexander Braverman, Weitong Zhang

arXiv 2609.35698首次发表:更新:

发表机构

HKUST; UNC Chapel Hill; Cornell University(香港科技大学; 北卡罗来纳大学教堂山分校; 康奈尔大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出双重正则化AIL算法,结合KL策略正则化与二次奖励惩罚,在有限时域MDP中证明其快速收敛率,并首次实现专家演示与在线交互的$\widetilde{O}(1/\epsilon)$样本复杂度。

AI 中文摘要

我们研究对抗模仿学习(AIL),其中智能体通过优化一个策略来学习模仿专家演示,该策略针对一个区分专家与学习者行为的对抗性奖励进行优化。历史上,奖励正则化和基于熵的策略正则化是诸如GAIL和LS-IQ等经验成功方法的关键组成部分,然而它们在有限样本下的优势仍未得到充分探索。我们在具有一般函数逼近的有限时域马尔可夫决策过程中,为联合正则化的AIL建立了快速收敛率。我们的无模型算法——双重正则化AIL(Dually Regularized AIL),将KL策略正则化与由专家和学习者占用度加权的二次奖励惩罚相结合。在K次在线回合和N条专家轨迹下,我们证明了固定正则化参数下正则化模仿差距的$\widetilde{O}\left(\frac{1}{K}+\frac{1}{N}\right)$界。我们的分析将针对一般凸奖励类别的在线镜像下降构造(用于控制来自有限专家数据和随机学习者反馈的估计误差)与对乐观KL正则化策略学习的精细分析相结合。据我们所知,Dually Regularized AIL是第一个在此正则化AIL目标下同时实现专家演示和在线交互的$\widetilde{O}\left(\frac{1}{\epsilon}\right)$样本复杂度的算法,即使面对随机专家也是如此。这些结果为AIL中奖励和策略正则化的互补统计优势提供了严格的刻画。

英文摘要

We study adversarial imitation learning (AIL), in which an agent learns to imitate expert demonstrations by optimizing a policy against an adversarial reward that distinguishes expert and learner behavior. Historically, reward regularization and entropy-based policy regularization are key components of empirically successful methods such as GAIL and LS-IQ, yet their finite-sample benefits remain underexplored. We establish fast rates for jointly regularized AIL in finite-horizon Markov decision processes with general function approximation. Our model-free algorithm, Dually Regularized AIL, combines KL policy regularization with a quadratic reward penalty weighted by expert and learner occupancies. With K online episodes and N expert trajectories, we prove a $\widetilde{O}\left(\frac{1}{K}+\frac{1}{N}\right)$ bound on the regularized imitation gap for fixed regularization parameters. Our analysis combines an online mirror descent construction for general convex reward classes to control estimation error from finite expert data and stochastic learner feedback, with a sharp analysis of optimistic KL-regularized policy learning. To the best of our knowledge, Dually Regularized AIL is the first algorithm to simultaneously achieve $\widetilde{O}\left(\frac{1}ε\right)$ sample complexity in both expert demonstrations and online interactions for this regularized AIL objective, even with stochastic experts. These results provide a rigorous characterization of the complementary statistical benefits of reward and policy regularization in AIL.

Comments33 pages, 1 table

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑