发表机构
Shenyang Institute of Automation, Chinese Academy of Sciences; University of Chinese Academy of Sciences; The University of Hong Kong; School of Computing and Data Science, The University of Hong Kong(中国科学院沈阳自动化研究所; 中国科学院大学; 香港大学; 香港大学计算与数据科学学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出离散扩散桥(DDB)框架,解决离散扩散在图像翻译与生成中的时空错位问题,通过混合吸收机制和信息引导噪声调度实现平衡,在多生成任务中表现优异且适配低采样步数场景。
AI 中文摘要
我们提出了离散扩散桥(Discrete Diffusion Bridges,简称DDB),这是一种旨在解决标准离散扩散在图像翻译与生成中存在的基础时空错位问题的新型框架。传统前向过程通过随机调度将数据损坏为纯掩码状态,引发双重错位:空间上,该纯掩码目标完全丢弃了源图像丰富的结构先验;时间上,随机掩码顺序与推理阶段采用的“易先难后”解码机制固有矛盾。为解决此问题,DDB在域间构建了直接且高效的轨迹。空间上,我们引入混合吸收机制,将吸收状态重新定义为掩码与源 token 的随机混合,有效将源先验作为空间锚点注入潜在空间;时间上,我们设计了信息引导的噪声调度,对语义变化进行量化,以在更早时间步优先损坏高信息区域,确保模型学习利用不变区域的鲁棒上下文解决困难的语义变化。大量实验验证了我们的框架在不同生成范式下的通用性与鲁棒性,DDB在文本引导的语义操作和纯结构图像翻译中均能有效平衡编辑对齐与结构保真度,同时天然补充文本到图像生成,并在极低采样步数下保证鲁棒的高质量解码。代码和模型可在该网址获取:this https URL。
英文摘要
We propose Discrete Diffusion Bridges (DDB), a novel framework designed to resolve the fundamental spatiotemporal misalignment of standard discrete diffusion in image translation and generation. By corrupting data into a pure mask state via a random schedule, the conventional forward process induces a twofold misalignment: spatially, this pure-mask destination entirely discards the rich structural priors of the source image; temporally, the random masking order inherently contradicts the ``easy-first, hard-last'' decoding mechanism used during inference. To address this, DDB constructs a direct and efficient trajectory between domains. Spatially, we introduce a hybrid absorption mechanism that redefines the absorbing state to a stochastic mixture of mask and source tokens, effectively injecting source prior as spatial anchors into the latent space. Temporally, we design an information-guided noise schedule that quantifies semantic variation to prioritize the corruption of high-information regions at earlier timesteps. This ensures the model learns to resolve difficult semantic changes using robust context from invariant regions. Extensive experiments validate the versatility and robustness of our framework across diverse generative paradigms. DDB effectively balances edit alignment with structural fidelity across both text-guided semantic manipulation and pure structural image translation, while inherently complementing text-to-image generation and guaranteeing robust high-quality decoding under extremely low sampling steps. Code and models are available at \href{https://github.com/HKU-HealthAI/DDB}{https://github.com/HKU-HealthAI/DDB}.
CommentsAccepted to ECCV 2026