AI 中文总结
研究图像到图像翻译中弱对齐数据问题,提出A²BM方法利用图像对对齐,在训练中纳入对齐分数区分真实对应,推理时控制翻译保真度,实验验证该方法比基线表现优,为弱对齐数据的图像翻译提供解决方案。
AI 中文摘要
配对的图像到图像翻译支撑着广泛的计算机视觉任务。桥接匹配和流匹配是强大框架,但标准公式假设训练对完美对齐。实际应用中常涉及弱对齐对。本文引入对齐感知桥接匹配(A²BM),在训练中利用图像对对齐,通过纳入对齐分数让模型区分真实语义对应。推理时用对齐分数控制翻译保真度。在合成实验和现实任务中验证,A²BM比基线 consistently improves translation fidelity over strong GAN-, diffusion-, and Schr{ö}dinger bridge-based baselines, establishing alignment conditioning as a principled solution for image translation models with weakly aligned data.
英文摘要
Paired image-to-image translation underpins a wide range of computer vision tasks, including image editing, sensor translation, and domain adaptation. Bridge matching and flow matching have recently emerged as powerful frameworks, extending diffusion models to arbitrary source and target distributions. However, their standard formulations assume perfectly aligned training pairs, treating all source-target correspondences as equally reliable. In practice, real-world applications often involve weakly aligned pairs due to changes of acquisition conditions, including e.g. asynchronous captures, different illuminations, or misregistration. In this work, we introduce Alignment-Aware Bridge Matching (A${}^2$BM), a bridge matching method that leverages image pairs alignment during training. By incorporating alignment scores, the model learns to disentangle true semantic correspondences from misalignment artifacts. At inference time, we use the alignment score as a control variable over translation fidelity, with strongly aligned outputs obtained when prompting the model with the highest alignment score. We validate A${}^2$BM on both controlled synthetic experiments and on challenging real-world tasks, including cross-sensor super-resolution and pixel-space unsupervised domain adaptation. In all settings, A${}^2$BM consistently improves translation fidelity over strong GAN-, diffusion-, and Schr{ö}dinger bridge-based baselines, establishing alignment conditioning as a principled solution for image translation models with weakly aligned data.