arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于高斯和Beta策略的基于图像的连续控制中的分歧正则化模仿学习

Disagreement-Regularized Imitation Learning for Image-Based Continuous Control with Gaussian and Beta Policies

Irving Giovani Bronzatti Petrazzini, Eric Aislan Antonelo

arXiv 2609.38407首次发表:更新:

发表机构

Federal University of Santa Catarina(圣卡塔琳娜联邦大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出分歧正则化模仿学习(DRIL),利用策略分歧作为奖励,在少演示图像连续控制中显著提升性能,Beta策略在更多演示时表现更优。

AI 中文摘要

目的:行为克隆在学习控制器访问演示分布之外的状态时可能会累积误差。本研究评估了分歧正则化模仿学习(DRIL),该方法将克隆策略之间的分歧转化为强化学习奖励,是否能改善基于图像的连续控制。方法:一项受控的CarRacing研究结合了高斯和Beta学习策略,演示来自裁剪高斯专家或内在有界Beta专家,一条或20条轨迹,确定性和随机评估,以及三个保留阶段:行为克隆、最高10集训练分数检查点,以及最终DRIL检查点。分歧集成在每个变体中包含五个高斯策略。每个保留策略在100个程序生成的剧集上评估。结果:分数选择的DRIL在少演示设置中产生了最大增益,在裁剪动作演示中比最强行为克隆均值提高61%,在有界动作演示中提高112%。在20条轨迹下,DRIL的优势缩小;在有界动作机制中,Beta行为克隆仍比最佳DRIL检查点高出约7%。实验还表明,分歧奖励的信息量随学习表示和训练阶段而变化。结论:DRIL可以显著改善少演示视觉连续控制,而有界Beta策略在更多演示可用时提供强大的行为克隆性能。结果强调了学习支持、集成响应和检查点选择的共同重要性。

英文摘要

Purpose: Behavior cloning can accumulate errors when a learned controller visits states outside the demonstrated distribution. This study evaluates whether Disagreement-Regularized Imitation Learning (DRIL), which converts disagreement among cloned policies into a reinforcement-learning reward, improves image-based continuous control. Methods: A controlled CarRacing study combines Gaussian and Beta learner policies, demonstrations from either a clipped Gaussian expert or an intrinsically bounded Beta expert, one or 20 trajectories, deterministic and stochastic evaluation, and three retained stages: behavior cloning, the highest 10-episode training-score checkpoint, and the final DRIL checkpoint. The disagreement ensemble contains five Gaussian policies in every variant. Each retained policy is evaluated over 100 procedurally generated episodes. Results: Score-selected DRIL produced its largest gains in the few-demonstration setting, improving over the strongest behavior-cloning mean by 61% with clipped-action demonstrations and by 112% with bounded-action demonstrations. With 20 trajectories, the advantage of DRIL narrowed; in the bounded-action regime, Beta behavior cloning remained about 7% above the best DRIL checkpoint. The experiments also show that the informativeness of the disagreement reward changes with the learner representation and training stage. Conclusion: DRIL can substantially improve few-demonstration visual continuous control, while bounded Beta policies provide strong behavior-cloning performance when more demonstrations are available. The results highlight the joint importance of learner support,ensemble response, and checkpoint selection.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑