arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.27885cs.LG

双向扩散桥用于多模态转换:往返之道

There and Back Again: Bidirectional Diffusion Bridges for Multimodality Translation

  • Stanford University(斯坦福大学)

机构由 AI 辅助整理,请以论文原文为准。

Gabe Guo, Elon Litman, Thanawat Sornwanee, Jose Blanchet, Stefano Ermon

AI总结:

针对现有多模态转换方法的不足,提出BIT双向图像-文本扩散桥,实现双向生成且性能优于基线。

AI中文摘要:

多模态转换(例如文本到图像)是生成式AI的核心任务。然而现有方法存在两点不足:一是遵循的生成路径未直接表示源模态,限制了部分采样算法的灵活性;二是单向性,无法实现逆转换(例如图像到文本)。我们提出BIT(双向图像-文本扩散桥),与现有方法不同,BIT直接从文本开始插值到图像,提供了源模态感知的生成路径,支持多样且灵活的采样算法;同时其端点条件化过程可实现从图像到文本的遍历,形成统一的双向生成框架。BIT通过随机微积分推导得出,生成易于模拟的SDE形式及可处理高维扩展的损失函数。实验表明,BIT与去噪扩散、确定性流基线具有竞争力,在多个视觉-语言及自然科学评估中表现优于这些基线。

英文摘要:

Multimodality translation (e.g., text-to-image) is a core generative AI task. However, existing approaches (1) follow generative paths that do not directly represent the source modality, limiting the flexibility of some sampling algorithms; and (2) are unidirectional, preventing inversion (e.g., image-to-text). We propose BIT: Bidirectional Image-Text Diffusion Bridges. In contrast to previous approaches, BIT starts directly from text and interpolates into images, providing (1) a source-aware generative path that enables diverse and flexible sampling algorithms; and (2) an endpoint-conditioned process that can be traversed from image to text, providing a unified, bidirectional generative framework. BIT is derived through stochastic calculus, yielding SDE forms amenable to simulation and tractable loss functions that scale to high dimensions. Our experiments show that BIT is competitive with denoising-diffusion and deterministic-flow baselines, and outperforms them on several vision--language and natural-science evaluations.

↑