arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

可扩展的多任务逆强化学习

Scalable Multi-Task Inverse Reinforcement Learning

Allen Tran, Jia Wan, Nathan Kallus, Aurélien Bibaut

arXiv 2610.00758首次发表:更新:

发表机构

Netflix; Massachusetts Institute of Technology; Cornell Tech(奈飞; 麻省理工学院; 康奈尔科技学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出一种低秩假设下的多任务逆强化学习方法,通过汇集多智能体数据缓解覆盖需求,实现可扩展的奖励恢复与策略迁移,并在实验中展现出更低遗憾和计算优势。

AI 中文摘要

通过学习可迁移的奖励函数,逆强化学习(IRL)能够在修改后的环境中对智能体进行反事实评估。这种迁移对覆盖范围提出了严格要求,因为目标环境会影响智能体的状态占用。我们提出了一种多任务IRL方法,在低秩假设下,将同一环境中具有不同奖励的多个智能体的数据汇集起来。除了缓解覆盖要求(即每个任务无需访问所有状态,只要其他任务访问过即可)之外,该方法还提供了在新环境下对多个任务进行可扩展评估的能力,因为计算密集型的规划随秩而非任务数量扩展。我们提供了在奖励恢复和新环境策略学习方面的有限样本保证。实验表明,我们的方法对有限覆盖具有鲁棒性,能够在每个任务的支持集内外恢复奖励,并以比基线更低的遗憾迁移到目标环境,且随着任务数量的增长,其相对于单任务方法的计算优势不断扩大。

英文摘要

By learning transferable rewards, inverse reinforcement learning (IRL) enables counterfactual evaluation of agents under modified environments. Such transfer places strict requirements on coverage since target environments affect agents' state occupancy. We propose a multi-task IRL method that pools data across multiple agents with different rewards in the same environment under a low-rank assumption. In addition to alleviating coverage requirements, so each task need not visit every state as long as others do, the method offers scalable evaluation of multiple tasks under new environments as computationally intensive planning scales with rank rather than the number of tasks. We provide finite sample guarantees on reward recovery and on policy learning in new environments. Experiments show our method is robust to limited coverage, recovers rewards on and off of each task's support, transfers to target environments at lower regret than baselines, with its computational advantage over per-task methods widening as tasks grow.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑