arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.20300cs.LG

通过边界感知数据增强提升离线强化学习的泛化性与鲁棒性

Improving Generalization and Robustness in Offline Reinforcement Learning via Boundary-Aware Data Augmentation

  • Tongji University(同济大学)

机构由 AI 辅助整理,请以论文原文为准。

Gong Gao, Weidong Zhao, Xianhui Liu

AI总结:

针对离线强化学习在分布内泛化与鲁棒性不足的问题,提出边界感知数据增强(BADA)方法,通过相邻状态构建插值边界生成合成数据,在多个基准上取得最优性能。

AI中文摘要:

当前的离线强化学习(ORL)算法往往过度拟合训练数据集,在部署到真实环境时表现出较差的分布内泛化性和鲁棒性,从而削弱了其有效性。现有方法通常利用计算机视觉中广泛使用的正则化技术来增强分布内泛化性和鲁棒性。然而,由于低级物理信号对分布偏移的高度敏感性,这些方法在分布内泛化性和鲁棒性方面仍存在显著局限,难以在复杂环境中实现稳定性能。为解决这一问题,我们从理论上分析了使用随机情节插值训练的行为策略和动作价值函数的误差界,揭示误差与状态间距离呈正相关。基于此见解,我们提出了一种名为边界感知数据增强(BADA)的方法,该方法利用相邻状态构建插值边界,从而生成更忠实保留原始数据分布的合成数据。我们首先在玩具环境中进行定性研究,表明BADA生成的混合样本在保持理想策略平滑性的同时,能准确重建多模态价值分布。在有限离线数据集上的大量实验进一步证明,BADA在多个基准上达到了最先进的性能。

英文摘要:

Current offline reinforcement learning (ORL) algorithms tend to overfit the training dataset and exhibit poor in-distribution generalization and robustness performance when deployed to real environments, thus compromising their effectiveness. Existing methods typically enhance in-distribution generalization and robustness by leveraging regularization techniques widely used in computer vision. However, due to the high sensitivity of low-level physical signals to distributional shifts, these methods still suffer from notable limitations in in-distribution generalization and robustness, making it difficult to achieve stable performance in complex environments. To address this issue, we theoretically analyze the error bounds of the behavior policy and action-value function trained with random episode interpolation, revealing that the error scales positively correlated with the distance between states. Based on this insight, we propose a method called $\bf{B}$oundary-$\bf{A}$ware $\bf{D}$ata $\bf{A}$ugmentation (BADA), which leverages neighboring states to construct interpolation boundaries, enabling the generation of synthetic data that more faithfully preserves the original data distribution. We first conduct qualitative studies in a toy environment, showing that BADA generates mixed samples that preserve desirable policy smoothness while accurately reconstructing multimodal value distributions. Extensive experiments on limited offline datasets further demonstrate that BADA attains state-of-the-art performance across diverse benchmarks.

↑