AI 中文总结
针对视觉强化学习中潜在距离选择影响性能的问题,提出PAMD(成对自适应马氏距离),它为测量潜在状态相似性参数化正定成对条件度量,是现有双模拟方法的简单插件,在视觉MuJoCo任务中显著提升了相关算法性能。
AI 中文摘要
许多视觉强化学习算法通过将潜在距离与由奖励和转移相似性诱导的行为距离相匹配来学习表示。在实践中,潜在距离的选择会强烈影响性能:使用固定的、预先指定的全局范数(如$\ell_p$范数或其他手工设计的度量)可能过于受限而无法捕捉行为距离。相比之下,无约束的成对距离可能会产生退化解,使度量损失降低但无法改善表示。为解决这一差距,我们引入了成对自适应马氏距离(PAMD),它参数化了一个正定的、成对条件度量来测量潜在状态相似性。PAMD是现有基于双模拟方法的简单插件,为固定的、预先指定的潜在距离提供了更具表现力且结构化的替代方案。我们在视觉MuJoCo连续控制任务上对方法进行了实证验证,当配备我们提出的距离时,几种基于双模拟的强化学习算法的最终性能得到了显著提升。
英文摘要
Many visual reinforcement learning (RL) algorithms learn representations by matching latent distances to a behavioral distance induced by reward and transition similarity. In practice, the choice of the latent distance can strongly affect performance: using a fixed, pre-specified global norms (e.g., $\ell_p$ norms or other hand-designed metrics) may be overly restrictive to capture the behavioral distance. In contrast, unconstrained pairwise distances may admit degenerate solutions that drive the metric loss down without improving the representation. To address this gap, we introduce **PAMD: Pairwise Adaptive Mahalanobis Distance**, which parameterizes a positive-definite, pair-conditioned metric for measuring latent state similarity. PAMD is a simple plug-in for existing bisimulation-based methods, offering a more expressive yet structured alternative to fixed, pre-specified latent distances. We empirically validate our method on visual MuJoCo continuous-control tasks, where final performance of several recent bisimulation-based RL algorithms is substantially improved when equipped with the distance we propose.
Comments9 pages, 6 figures, plus appendix. Accepted to the 43rd International Conference on Machine Learning (ICML 2026)