GFlowNets中的信息几何前向策略训练
Information-Geometric Forward Policy Training in GFlowNets
查看机构详情
- University of Nottingham(诺丁汉大学)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
本研究基于GFlowNets的信息几何构建前向策略训练,推导轨迹Fisher分解,提出三种计算场景及图模型替代方案,经实证验证其可实现结构感知的前向策略训练。
中文摘要 AI 辅助
生成流网络(GFlowNets)已成为一种灵活的框架,用于对离散及混合离散-连续对象进行摊销推断,仅需通过奖励指定的未归一化目标密度。在本研究中,我们通过诱导轨迹采样器的信息几何,构建GFlowNets中的前向策略训练。将前向策略视为诱导轨迹采样器,我们表明其固有一阶几何由轨迹族的Fisher-Rao度量给出,且当对应的Fisher信息可计算或可准确近似时,相关的自然梯度提供了规范的局部更新。我们推导了轨迹Fisher的精确分解,将其分解为每步条件二阶矩,这阐明了在共享参数化下时间评分相互作用何时消失、何时保持密集耦合。这产生三种计算场景:具有易处理精确Fisher信息的场景、预期Fisher的蒙特卡洛估计量已足够的场景,以及可利用结构的场景,其中目标局部性或因式分解能产生Fisher预期的准确近似。在后一场景中,精确边缘化、分隔符方法和信念传播等图模型工具为自然梯度更新提供了原则性替代方案。所得框架将目标结构转化为优化几何,为GFlowNets中结构感知前向策略训练提供了一条易处理的途径。我们通过比较黎曼优化与欧氏优化下收敛性及探索行为的示例,对该框架进行了实证说明。
英文摘要
Generative Flow Networks (GFlowNets) have emerged as a flexible framework for amortised inference over discrete and mixed discrete-continuous objects, requiring only an unnormalised target density specified through a reward. In this work, we formulate forward-policy training in GFlowNets through the information geometry of the induced trajectory sampler. Treating the forward policy as an induced trajectory sampler, we show that its intrinsic first-order geometry is given by the Fisher-Rao metric of the trajectory family, and that the associated natural gradient provides the canonical local update whenever the corresponding Fisher information is computable or accurately approximable. We derive an exact decomposition of the trajectory Fisher into per-step conditional second moments, which clarifies when temporal score interactions vanish and when dense couplings remain under shared parameterisation. This leads to three computational regimes: settings with tractable exact Fisher information, settings where Monte Carlo estimators of the expected Fisher are sufficient, and structure-exploitable settings in which target locality or factorisation yields accurate approximations of the Fisher expectation. In the latter case, graphical-model tools such as exact marginalisation, separator methods, and belief propagation provide principled surrogates for natural-gradient updates. The resulting framework turns target structure into optimisation geometry and yields a tractable route to structure-aware forward-policy training in GFlowNets. We illustrate the framework empirically through examples comparing convergence and exploration behaviour under Riemannian and Euclidean optimisation.