arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.10634cs.LG

IADD-TR:结合干预感知动力学解耦与定向正则化的基于模型强化学习方法

IADD-TR: Intervention-Aware Dynamics Decoupling with Targeted Regularization for Model-Based Reinforcement Learning

Zefeng Liang, Jie Qiao, Ruichu Cai, Weilin Chen, Zhifeng Hao

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出IADD-TR框架,通过干预感知动力学解耦与定向正则化解决MBRL中策略诱导的数据偏差问题,经MuJoCo任务实验验证可提升样本效率并取得有竞争力回报。

中文摘要 AI 辅助

基于模型的强化学习(MBRL)通过学习环境动力学生成合成经验,是一种有前景的样本高效决策方法。目前已有大量方法通过不确定性估计、模型正则化和保守值学习来改进MBRL的动力学预测与策略优化,但这些方法通常将转移模型与评判器视为整体预测器,忽略了策略诱导的数据偏差,导致动作与环境演化纠缠,且动作覆盖不均可能扭曲策略改进所用的反事实价值估计。为解决该问题,本文提出IADD-TR,这是结合干预感知动力学解耦(IADD)与定向正则化(TR)的统一框架。IADD将转移分解为动作干预阶段与无动作自然演化阶段,利用零动作锚点解决该两阶段分解的非唯一性问题以实现鲁棒泛化,其潜在分量与状态对齐分量分别在可逆块内变换与逐点意义上可识别。对于策略学习,本文从重放状态策略梯度泛函的有效影响函数中推导TR,TR为评判器增加动作密度缩放的残差校正项并优化定向损失,当评判器或重放动作密度被一致指定时,可产生双重鲁棒的策略梯度估计。在五项MuJoCo任务上的大量实验表明,IADD-TR在取得有竞争力回报的同时提升了样本效率。

英文摘要

Model-based reinforcement learning (MBRL), which learns environment dynamics to generate synthetic experience, is a promising approach to sample-efficient decision making. Numerous methods have been developed to improve dynamics prediction and policy optimization for MBRL through uncertainty estimation, model regularization, and conservative value learning. However, these methods typically treat the transition model and critic as monolithic predictors, overlooking the policy-induced data bias. Consequently, action can become entangled with environmental evolution, while uneven action coverage may distort the counterfactual value estimates used for policy improvement. To address this, we propose IADD-TR, a unified framework combining Intervention-Aware Dynamics Decoupling (IADD) and Targeted Regularization (TR). IADD factorizes transitions into an action-intervention stage and an action-free natural evolution stage, using a zero-action anchor to resolve the non-uniqueness of this two-stage factorization for robust generalization. Its latent and state-aligned components are identifiable up to an invertible within-block transformation and pointwise, respectively. For policy learning, we derive TR from the efficient influence function of a replay-state policy-gradient functional. TR augments the critic with an action-density-scaled residual correction and optimizes a targeted loss, yielding doubly robust policy-gradient estimation when either the critic or the replay action density is consistently specified. Extensive experiments on five MuJoCo tasks show that IADD-TR achieves competitive returns with improved sample efficiency.

发表机构

  • School of Computer Science, Guangdong University of Technology(广东工业大学计算机学院)
  • College of Science, Shantou University(汕头大学理学院)

机构由 AI 辅助整理,请以论文原文为准。

↑