arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.29617cs.LGcs.AIstat.ML

基于值函数的模仿学习中,何时在线交互是有帮助的?表示权衡研究

When Does On-Policy Interaction Help? Representational Tradeoffs in Value-Based Imitation Learning

Luca Viano, Antoine Moulin, Audrey Huang, Volkan Cevher, Philip Amortila, Dylan J. Foster

首次发表
浏览论文内容

中文总结 AI 辅助

本文针对基于值函数的模仿学习,提出交互式在线算法OVI,发现专家交互可降低学习者的表示要求,且OVI在多数场景下性能优于现有方法。

中文摘要 AI 辅助

模仿学习(Imitation Learning, IL)用于训练智能体从演示中复刻专家行为,支撑着从机器人学到语言模型训练的各类应用。标准方法如行为克隆(Behavior Cloning, BC)存在误差累积和性能停滞问题,尤其当学习者无法完美表示专家策略时(如蒸馏场景中常见情况)。两种干预手段经实证被广泛认为可提升性能:沿学习者自身轨迹交互式查询专家,以及在生成策略时利用途中的值函数估计而非直接拟合专家的完整动作分布。本研究探究这些改进的本质及其可能出人意料的相互作用。核心发现为:专家交互可降低学习者的表示要求,仅需模型能实现专家的值函数,无需满足更严格的实现专家策略本身的要求。具体而言,本文提出OVI算法,这是一种交互式在线模仿学习算法,只要学习者能表示专家的值函数,该算法就具有统计效率;若能使用线性最大化神谕,则具有计算效率。同时补充了负结果,证明交互是必要的:仅在专家值可实现性之外无更强假设时,任何离线模仿学习算法的复杂度都必须与专家策略类的复杂度成比例。实证结果验证了上述发现:OVI的性能优于离线策略型方法(BC)、交互式策略型方法(DAgger)及离线值型模仿学习方法,且当学习者网络的表达能力远弱于专家时,性能提升最为显著。

英文摘要

Imitation learning (IL)---training an agent to replicate expert behavior from demonstrations---underpins applications from robotics to language model training. Standard approaches such as Behavior Cloning (BC) are known to suffer from compounding errors and performance plateaus, particularly when the learner cannot perfectly represent the expert's policy (as is typical, e.g., in distillation). Two interventions are widely understood empirically to improve performance: querying the expert interactively along the learner's own trajectories, and using value function estimation en route to generating a policy rather than directly fitting the expert's full action distribution. We investigate the nature of these improvements and their potentially surprising interplay. Our main finding is that expert interaction relaxes the representational demands on the learner: one only needs a model capable of realizing the expert's value function, bypassing the (often stricter) requirement of realizing the expert's policy itself. Concretely, we introduce OVI, an interactive on-policy IL algorithm that is statistically efficient whenever the learner can represent the expert's value function and computationally efficient given access to a linear maximization oracle. We complement this with a negative result showing that interaction is necessary. Namely, without stronger assumptions beyond expert-value realizability alone, any offline IL algorithm must scale with the complexity of the expert policy class. Our findings bear out empirically. OVI outperforms offline policy-based (BC), interactive policy-based (DAgger), and offline value-based IL methods, with the largest gains when the learner network is substantially less expressive than the expert's.

↑