泛化与引导:用于少样本逆强化学习的奖励分解
Generalize and Guide: Decomposing Rewards for Few-Shot Inverse Reinforcement Learning
浏览论文内容
中文总结 AI 辅助
研究少样本逆强化学习(FM - IRL)问题,提出多任务判别器近邻引导的IRL(MPG)方法,通过学习两个互补奖励组件,在多种有显著变化的任务上验证有效性,平均成功率达81.2%,优于基线。
中文摘要 AI 辅助
逆强化学习(IRL)为从演示中学习提供了一个强大的框架。然而,现实世界的任务通常存在很大的自然变化(例如,拿起不同形状的杯子),使得收集能在每种可能场景下完全指定新任务的演示变得不切实际。实际上,虽然目标任务的演示有限,但获取异构但相关行为的数据集往往更容易。这就引发了多任务演示的少样本IRL(FM - IRL)问题,即智能体必须仅从有限数量的目标任务演示以及相关任务的充分演示和在线智能体经验中学习具有大量变化的新任务。为此,我们必须既恢复新任务的专家分布,又在智能体偏离时提供指导。我们引入了多任务判别器近邻引导的IRL(MPG),它学习两个互补的奖励组件:(1)一个可泛化的判别器,跨相关任务传递共享结构以识别新任务中的专家行为;(2)一个近邻函数,测量状态与专家行为的偏离程度并在探索过程中提供纠正指导。我们在多种具有显著变化(如物体配置、桌子布局和初始机器人姿势)的具有挑战性的导航和操纵任务上证明了我们方法的有效性,平均成功率达到81.2%,比最强的每个任务基线平均高出24.7个百分点。
英文摘要
Inverse reinforcement learning (IRL) provides a powerful framework for learning from demonstrations. However, real-world tasks often exhibit substantial natural variations (e.g., picking up mugs with varying shapes), making it impractical to collect demonstrations that fully specify a new task under every possible scenario. In practice, while demonstrations for the target task are limited, it is often easier to obtain datasets of heterogeneous but related behaviors. This motivates the problem of few-shot IRL with multi-task demonstrations (FM-IRL), where an agent must learn a new task with substantial variations from only a limited number of target-task demonstrations, together with sufficient demonstrations of related tasks and online agent experience. To do so, we must both recover the expert distribution of the new task and provide guidance when the agent deviates from it. We introduce Multitask discriminator Proximity-Guided IRL (MPG), which learns two complementary reward components: (1) a generalizable discriminator that transfers shared structure across related tasks to identify expert behavior in a new task, and (2) a proximity function that measures how far a state deviates from expert behavior and provides corrective guidance during exploration. We demonstrate the effectiveness of our method on multiple challenging navigation and manipulation tasks under significant variations (e.g., object configurations, table layouts, and initial robot poses), achieving an average success rate of 81.2%, outperforming the strongest per-task baseline by an average of 24.7 percentage points.