arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.16090cs.LGcs.AI

DADiff:用于强化学习的扩散驱动跨域策略适应

DADiff: Diffusion-Driven Cross-Domain Policy Adaptation for Reinforcement Learning

  • Arizona State University(亚利桑那州立大学)
  • University of Maryland, College Park(马里兰大学帕克分校)

机构由 AI 辅助整理,请以论文原文为准。

Hanyang Chen, Anirudh Satheesh, Longchao Da, Hua Wei

AI总结:

研究强化学习中跨域策略适应问题,提出基于扩散的DADiff框架,利用源域和目标域生成轨迹差异估计动态不匹配,开发奖励修改和数据选择变体,实验表明该方法性能优于现有方法,有效解决动态不匹配。

AI中文摘要:

在强化学习中,跨域转移策略面临着重大挑战,因为源域和目标域之间存在动态不匹配。本文考虑在线动态适应设置,即在源域使用足够数据训练策略,同时仅允许与目标域进行有限交互。现有一些工作通过使用域分类器、值引导数据过滤或表示学习来解决动态不匹配问题。相反,我们从生成建模角度研究域适应问题。具体来说,我们引入DADiff,这是一个基于扩散的框架,它利用下一状态生成过程中源域和目标域生成轨迹之间的差异来估计动态不匹配。我们开发了奖励修改和数据选择变体来使策略适应目标域。我们还进行了理论分析,表明给定策略在两个域之间的性能差异受生成轨迹偏差的限制。我们在具有各种转移的环境中进行了广泛实验,结果表明我们的方法比现有方法具有更好的性能,有效解决了动态不匹配问题。

英文摘要:

Transferring policies across domains poses a vital challenge in reinforcement learning, due to the dynamics mismatch between the source and target domains. In this paper, we consider the setting of online dynamics adaptation, where policies are trained in the source domain with sufficient data, while only limited interactions with the target domain are allowed. There are a few existing works that address the dynamics mismatch by employing domain classifiers, value-guided data filtering, or representation learning. Instead, we study the domain adaptation problem from a generative modeling perspective. Specifically, we introduce DADiff, a diffusion-based framework that leverages the discrepancy between source and target domain generative trajectories in the generation process of the next state to estimate the dynamics mismatch. Both reward modification and data selection variants are developed to adapt the policy to the target domain. We also provide a theoretical analysis to show that the performance difference of a given policy between the two domains is bounded by the generative trajectory deviation. More discussions on the applicability of the variants and the connection between our theoretical analysis and the prior work are further provided. We conduct extensive experiments in environments with various shifts to validate the effectiveness of our method. The results demonstrate that our method provides superior performance compared to existing approaches, effectively addressing the dynamics mismatch. We provide the code of our method at https://github.com/hanyang-chen/DADiff-release

补充信息

↑