arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

VIGOR:基于模型强化学习中潜在空间一致性的零样本视觉泛化

VIGOR: Zero-Shot Visual Generalization via Latent-Space Consistency in Model-Based Reinforcement Learning

Mingyu Park, Samyeul Noh, Hyun Myung, Donghwan Lee

arXiv 2610.02801首次发表:更新:

发表机构

KAIST; ETRI(韩国科学技术院; 韩国电子通信研究院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

VIGOR通过潜在空间一致性实现基于模型强化学习对未见视觉干扰的零样本泛化,在DMC和Robosuite上分别超越次优基线3.4%和43.6%。

AI 中文摘要

基于模型的强化学习(MBRL)通过在学习的潜在动态中进行规划实现了强大的样本效率,但在背景变化、光照变化或相机移动等未见过的视觉干扰下,其性能会大幅下降。与无模型强化学习不同(在无模型强化学习中,编码器扰动仅影响单步预测),MBRL存在两级脆弱性:视觉干扰首先将编码器输出推离分布,然后这些误差在规划范围内通过递归潜在展开而复合。我们提出了基于模型的强化学习中的视觉泛化潜在空间一致性(VIGOR),该框架能够实现对未见过的视觉干扰的零样本泛化,同时保留其MBRL骨干的样本效率。VIGOR集成了三个相互依赖的组件:(i)非对称弱到强增强,在单个批次内配对仅弱增强和弱到强增强的潜在视图;(ii)动态级一致性,通过直接潜在回归强制执行增强不变的转移预测;(iii)编码器级稳定化,在动态级一致性施加的跨增强监督下防止编码器漂移。在DeepMind控制套件(DMC)和Robosuite上的评估表明,VIGOR优于最先进的无模型和基于模型的基线,在DMC上超过第二好的基线3.4%,在Robosuite上超过43.6%。消融研究进一步表明,VIGOR的鲁棒性与增强无关:将默认增强替换为来自不同扰动族的替代方案仍能保持强大的泛化能力,确认了潜在空间一致性而非增强选择驱动了鲁棒性。

英文摘要

Model-based reinforcement learning (MBRL) achieves strong sample efficiency by planning within learned latent dynamics, yet its performance degrades substantially under unseen visual distractions such as background variations, lighting changes, or camera shifts. Unlike model-free RL, where encoder perturbations affect only single-step predictions, MBRL suffers from a two-level vulnerability: visual distractions first push encoder outputs out of distribution, and these errors then compound through recursive latent rollouts over the planning horizon. We propose visual generalization via latent-space consistency in model-based RL (VIGOR), a framework that enables zero-shot generalization to unseen visual distractions while retaining the sample efficiency of its MBRL backbone. VIGOR integrates three interdependent components: (i) asymmetric weak-to-strong augmentation, which pairs weak-only and weak-to-strong latent views within a single batch; (ii) dynamics-level consistency, which enforces augmentation-invariant transition predictions through direct latent regression; and (iii) encoder-level stabilization, which prevents encoder drift under the cross-augmentation supervision imposed by dynamics-level consistency. Evaluations on the DeepMind Control Suite (DMC) and Robosuite show that VIGOR outperforms state-of-the-art model-free and model-based baselines, surpassing the second-best baseline by 3.4% on DMC and 43.6% on Robosuite. Ablations further show that VIGOR's robustness is augmentation-agnostic: replacing the default augmentation with alternatives from distinct perturbation families preserves strong generalization, confirming that latent-space consistency, not the augmentation choice, drives robustness.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑