发表机构
Institute of Automation, Chinese Academy of Sciences; The University of Manchester; Capital University of Physical Education And Sports; Shanghai Jiao Tong University(中国科学院自动化研究所; 曼彻斯特大学; 首都体育学院; 上海交通大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对SAR到光学图像翻译中扩散模型推理慢、细节保真度低的问题,提出基于流匹配与对比学习的一步式模型ContraFM-S2O,在SAR2Opt和QXS数据集上达到最优并降低推理延迟。
AI 中文摘要
近年来,扩散模型和基于GAN的模型凭借高质量生成和稳定训练等优势,已成为SAR到光学图像翻译的主流方法。然而,它们存在推理延迟高以及生成的光学图像细节保真度低的问题,常导致边缘模糊和精细纹理丢失。为此,我们提出了ContraFM-S2O,一种基于流匹配的SAR到光学图像翻译模型。与传统扩散模型不同,ContraFM-S2O在训练中学习预测速度场,并在推理时求解ODE而非SDE,以提高采样效率。此外,ContraFM-S2O用沿插值路径的平均速度替代瞬时速度,实现一步式SAR到光学图像翻译,并利用对比学习提升生成光学图像的质量。实验表明,我们的模型在SAR2Opt和QXS数据集上达到了最先进水平,优于基线模型,并通过一步生成降低了推理延迟。
英文摘要
In recent years, diffusion models and GAN-based models have become the mainstream approaches for SAR-to-optical image translation, owing to their advantages, such as high-quality generation and stable training. However, they have shortcomings such as high inference latency and the generated optical images suffer from low detail fidelity, often resulting in blurred edges and loss of fine textures. Thus, we propose ContraFM-S2O, which is a flow matching-based model for SAR-to-optical image translation. Unlike conventional diffusion models, ContraFM-S2O learns to predict the velocity field in training and solves ODE instead of SDE during inference to improve the sampling efficiency. In addition, ContraFM-S2O replaces instantaneous velocity with average velocity along the interpolation path to realize one-step SAR-to-optical image translation and uses contrastive learning to improve the quality of the generated optical images. Experiments show our model achieves state-of-the-art on SAR2Opt and QXS datasets, outperforming baselines, and reduces inference latency via one-step generation.