AI 中文总结
该研究揭示跨维推断中学习型提议在平衡态最优,通过实验验证其加速模型阶数混合的效果,开发了开源工具HyperWave,可跨领域应用于源计数与信号重建。
AI 中文摘要
从图像中源的计数到混合模型,在推断模型维度(解释数据所需的分量数量)的同时推断参数是一个普遍存在的问题,可逆跳转马尔可夫链蒙特卡罗(reversible-jump Markov chain Monte Carlo)可精确求解该问题,但混合速度较慢。学习型提议在固定维度场景中已得到广泛应用,但其能否加速维度变换移动本身在很大程度上仍未得到验证。我们表明,答案具有结构性起源:维度变换的出生移动(birth move)的最优提议在运行的不同阶段是不同的对象。在拟合模型的构建过程中,它必须匹配当前残差——这是一个状态依赖量,无状态依赖的网络无法表示;但在平衡态时,它退化为后验分布的单分量边际,而这正是自适应归一化流(adaptive normalizing flow)从采样器自身历史中学习的分布。因此,无状态依赖的学习型出生提议在一个阶段无用,而在另一阶段最优。受控实验证实了这一归因:应用精确的Metropolis--Hastings校正(该校正对任意网络均保持目标分布不变)时,学习型出生移动虽未改变接受率,但加速了模型阶数混合——在十次种子基准测试中,它在十次运行中的六次达到了预设的停止规则,通常速度快数倍,而一个经过强调优的手动基线仅在一次运行中达到该规则(单侧p值为0.03);一项隔离实验显示,在模型内部部署相同的流无任何增益。该方法未做特定领域假设,可在科学领域(包括地面和空间探测器的引力波、头皮脑电图记录)的含噪图像中计数源并重建信号,我们将该方法作为开源包HyperWave发布。
英文摘要
Inferring the dimension of a model - the number of components needed to explain data - jointly with the parameters is a pervasive problem, from counting sources in an image to mixture modeling, and reversible-jump Markov chain Monte Carlo solves it exactly but mixes slowly. Learned proposals are well established at fixed dimension, but whether they can accelerate the dimension-changing moves themselves has remained largely untested. We show that the answer has a structural origin: the optimal proposal for the dimension-changing birth move is a different object in different phases of the run. While the fit is being assembled it must match the current residual - a state-dependent quantity no state-independent network can represent - but at equilibrium it degenerates to the posterior's single-component marginal, which is exactly the distribution an adaptive normalizing flow learns from the sampler's own history. A learned state-independent birth proposal is therefore useless in one phase and optimal in the other. Controlled experiments confirm the attribution: applied with an exact Metropolis--Hastings correction that leaves the target invariant for any network, the learned births leave acceptance rates unchanged yet accelerate model-order mixing - in a ten-seed benchmark they meet a pre-specified stopping rule in six of ten runs, typically several times sooner, where a strong hand-tuned baseline meets it in one (one-sided p=0.03) - and an isolation experiment shows the same flow deployed within-model buys nothing. Making no domain-specific assumptions, the same sampler counts sources in a noisy image and reconstructs signals across scientific domains, including gravitational waves from ground- and space-based detectors and a scalp EEG recording. We release the method as HyperWave, an open-source package.
Comments23 pages, 5 figures, 2 tables