发表机构
Macau University of Science and Technology; Pengcheng Laboratory; Sun Yat-sen University; Institute of Automation, Chinese Academy of Sciences (CASIA)(澳门科技大学; 鹏城实验室; 中山大学; 中国科学院自动化研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对二值扩散模型低NFE采样时样本质量下降的问题,提出伯努利流模型,通过定义连续全局伯努利概率流实现自洽低NFE采样,在LSUN Churches数据集上仅用16步采样即达到9.22的FID,性能优于离散基线。
AI 中文摘要
二值扩散模型通常需要大量函数评估次数(NFEs)才能生成高质量样本,导致实际推理的计算成本高昂。在不使用蒸馏或额外训练的前提下减少NFEs同时保持样本质量,仍是一项重大挑战。现有二值扩散模型定义了离散的单步前向路径,随后推导反向后验分布;在需要跨步采样的低NFE场景中,它们用单步似然转移近似真实多步似然,这会严重降低样本质量。为解决这一根本局限并将生成动力学与固定离散时间步解耦,我们提出伯努利流模型(BFM)。BFM不依赖顺序单步马尔可夫扩散链,而是在数据分布与纯噪声之间定义统一的连续全局伯努利概率流路径,由此推导任意时间区间上的解析闭式后验转移。因此,减少推理NFEs不再是跳过离散步的近似操作,仅需在新时间网格上重新计算解析后验即可。这消除了离散链固有的训练-推理结构不匹配,实现自洽的低NFE采样。实验表明,BFM对激进的NFE减少具有极强鲁棒性:在LSUN Churches 256x256数据集上,采用256步训练的BFM仅用16步采样即可达到9.22的FID,而最先进的离散基线则降至204.10;在标准全步推理下,BFM也与连续和离散生成基线保持竞争力。这些结果确立BFM为一种理论严谨、自洽且实用有效的快速二值数据生成框架。
英文摘要
Binary diffusion models typically require a large number of function evaluations (NFEs) to generate high-quality samples, making practical inference computationally expensive. Reducing NFEs while preserving sample quality without distillation or additional training remains a significant challenge. Existing binary diffusion models define a discrete one-step forward path and then derive the reverse posterior. In low-NFE settings requiring cross-step sampling, they approximate the true multi-step likelihood with a single-step likelihood transition, which severely degrades sample quality. To address this fundamental limitation and decouple the generative dynamics from fixed discrete time steps, we propose Bernoulli Flow Models (BFM). Rather than relying on sequential one-step Markov diffusion chains, BFM defines a unified continuous global Bernoulli probability flow path between data distributions and pure noise, from which we derive analytical closed-form posterior transitions over arbitrary time intervals. Consequently, reducing the inference NFE is no longer an approximation based on skipping discrete steps; it only requires re-evaluating the analytical posterior over a new time grid. This eliminates the structural training-inference mismatch inherent to discrete chains and yields self-consistent low-NFE sampling. Experiments show that BFM is highly robust to aggressive NFE reduction. On LSUN Churches 256x256, a BFM trained with 256 steps achieves an FID of 9.22 using only 16 sampling steps, whereas the state-of-the-art discrete baseline degrades to 204.10. BFM also remains competitive with continuous and discrete generative baselines under standard full-step inference. These results establish BFM as a theoretically rigorous, self-consistent, and practically effective framework for fast binary data generation.