arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.08644cs.LG

依赖路径的离散摊销推理

Path-dependent Discrete Amortized Inference

Tiago da Silva, Esmeralda S. Whitammer, Salem Lahlou

首次发表
浏览论文内容

中文总结 AI 辅助

针对从未归一化后验采样组合离散对象的问题,提出依赖路径的离散摊销推理方法,通过可学习潜动力学系统提升MDP,解决马尔可夫假设的缺陷,实验显示其学习收敛更快、状态空间探索更优。

中文摘要 AI 辅助

我们考虑从给定的未归一化后验分布中采样组合式离散对象的问题。值得注意的是,近期研究表明,通过学习确定性马尔可夫决策过程(MDP),按后验比例逐步构建每个对象,可高效解决该问题。然而,本研究证实,马尔可夫假设既会阻碍训练期间的信号传播,又会因状态混叠而灾难性降低学习到的采样器的表达能力。为解决这些问题,我们提出用可学习的潜动力学系统提升MDP,使基础策略能依赖整个过去轨迹,而非仅当前状态。据此,我们将所提方法称为依赖路径的离散摊销推理。重要的是,我们可证明,该方法将现有离散摊销采样器的学习算法扩展到了我们的设置中。在标准基准问题的实验中,我们还表明,与现有技术相比,本方法常能实现更快的学习收敛和更优的状态空间探索。

英文摘要

We consider the problem of sampling compositional and discrete objects from a given unnormalized posterior distribution. Notably, recent studies have shown that this problem can be efficiently solved by learning a deterministic Markov Decision Process (MDP) that progressively builds each object in proportion to the posterior. In this work, however, we demonstrate that the Markovian assumption can both hamper signal propagation during training and catastrophically reduce the learned sampler's expressivity due to state aliasing. To address these issues, we propose lifting the MDP with a learnable latent dynamical system that allows the underlying policy to depend on the entire past trajectory---and not only on the current state. In view of this, we refer to the resulting method as path-dependent discrete amortized inference. Importantly, we provably extend existing learning algorithms for discrete amortized samplers to our setting. In experiments on standard benchmark problems, we also show that our approach often leads to faster learning convergence and improved state space exploration relatively to prior techniques.

补充信息

↑