区分双模拟度量:通过双因果最优传输进行参数马尔可夫链拟合的框架
Differentiating Bisimulation Metrics: A Framework for Parametric Markov Chain Fitting via Bicausal Optimal Transport
- Universitat Pompeu Fabra(庞培法布拉大学)
- University of Texas at Austin(德克萨斯大学奥斯汀分校)
机构由 AI 辅助整理,请以论文原文为准。
中文总结 AI 辅助
本文提出可微双因果最优传输(D-BOT)算法,通过将双模拟度量作为可微线性规划,利用包络定理求闭式梯度,交替优化以拟合参数马尔可夫链,适用于状态压缩、模型学习和模仿学习。
中文摘要 AI 辅助
顺序决策中的许多问题,例如从观测进行模仿学习、状态空间压缩、世界模型学习以及模拟到现实迁移,都可以归结为学习一个模型,使得与目标过程之间的距离度量最小化。我们考虑这一通用框架,并将双模拟度量(等价于双因果最优传输,BOT)作为要最小化的距离度量。我们证明,由于BOT可以表述为线性规划(LP),因此它相对于模型动态是可微的。然后,我们通过包络定理应用于LP鞍点,推导出精确的闭式梯度。结果是一个通用算法——可微双因果最优传输(D-BOT),可应用于上述每个问题。所提出的算法通过交替进行距离计算和梯度步骤来学习最佳模型。我们将D-BOT应用于三种不同设置:状态空间压缩、参数模型学习和从观测进行模仿学习(ILfO)。我们展示的实验结果证实了这三种实例化的可行性。
英文摘要
Many problems in sequential decision-making, such as imitation learning from observations, state-space compression, world-model learning, and sim-to-real transfer, can be reduced to learning a model such that a notion of distance with respect to the target process is minimized. We consider this general framework and consider the bisimulation metric, equivalently Bicausal Optimal Transport (BOT), as the notion of distance to minimize. We show that BOT, since it can be formulated as a linear program (LP), is differentiable with respect to the model dynamics. We then derive an exact closed-form gradient via the envelope theorem applied to the LP saddle point. The result is a general algorithm, Differentiable Bicausal Optimal Transport (D-BOT), that can be applied to each of the problems above. The proposed algorithm learns the best model by alternating between distance computation and gradient steps. We apply D-BOT for three different settings: state-space compression, parametric model learning, and imitation learning from observations (ILfO). We show empirical results that confirm the viability of all three instantiations.