arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

用于生成式 3D 先验的投影能量匹配

Projected Energy Matching for Generative 3D Priors

Daniel Barco, Michal Balcerak, Suprosanna Shit, Chinmay Prabhakar, Philipp Denzel, Bjoern Menze, Frank-Peter Schilling

arXiv 2607.07749首次发表:更新:

发表机构

Zurich University of Applied Sciences (ZHAW); University of Zurich (UZH)(苏黎世应用科学大学; 苏黎世大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对能量匹配在高维 3D 数据训练的挑战,提出投影能量匹配框架,引入亥姆霍兹提炼和负缓存策略,用于医学 CT 逆问题如稀疏视图重建,减少计算量并实现高保真重建。

AI 中文摘要

能量匹配已成为一个强大的生成框架,它通过单个与时间无关的标量势将流模型效率与基于能量的模型(EBM)的显式似然性结合起来。然而,直接在高维 3D 数据上训练这个势在计算上仍然具有挑战性。虽然提炼预训练流模型可规避一些初始训练成本,但速度场不可避免地包含非保守旋转伪影(旋度)。强制严格保守的标量势与这个无约束场匹配会产生“结构冲突”,降低生成质量和模式覆盖。为解决此问题,提出投影能量匹配,它引入亥姆霍兹提炼以利用哈钦森迹估计器将旋转噪声明确吸收到辅助残差网络,随后用负缓存细化这种情况,负缓存是一种内存高效策略,跨微批次重用负样本,使对比训练中带梯度累积的采样易于处理。将该方法部署为真实世界医学 CT 逆问题(特别是稀疏视图重建)的无条件先验。最终,摊销管道将总计算量减少到标准能量匹配所需的一小部分,同时实现高保真重建并成功解决严重测量伪影。

英文摘要

Transport-based generative models, which learn a time-dependent vector field that moves noise to data, have become a dominant paradigm. However, these models typically do not explicitly encode the data distribution. Energy-based models (EBMs) instead represent the data distribution explicitly through a scalar energy landscape, which Energy Matching learns by combining transport learning with contrastive refinement. Its transport objective, however, fits energy gradients to stochastic targets, whose variance degrades the training signal at scale. We introduce Projected Energy Matching, which learns this landscape through a more stable route: we first train a time-independent transport teacher, then freeze it and fit the negative energy gradient to its predicted velocities. This projection replaces noisy transport targets with deterministic supervision, while contrastive refinement shapes the landscape near the data manifold. On CIFAR-10, gradient-noise analysis reveals a cleaner training signal, accompanied by faster convergence than Energy Matching at matched, teacher-free training budgets. In the latent space of CT volumes, our method enables 3D CT generation and achieves better FID scores than flow models. The learned scalar potential serves as a zero-shot prior for the ill-posed inverse problem of sparse-view cone-beam CT reconstruction. By making explicit energy landscapes practical at volumetric scale, this work opens a path to wider adoption of energy-based formulations, bringing their flexibility to high-dimensional generation and inverse problems.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑