发表机构
Korea Advanced Institute of Science and Technology (KAIST)(韩国科学技术院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
GeoSET是首个SAR到EO图像翻译的全能基础模型,通过预训练父模型和LoRA适应,仅更新0.60%参数,在六个基准上取得SOTA结果。
AI 中文摘要
成对的合成孔径雷达(SAR)和光电(EO)图像在传感器、分辨率和地理区域方面日益丰富。然而,现有的SAR到EO图像翻译(SET)方法通常仅在单一、有限规模的数据集上训练,产生针对特定传感条件的专用模型。我们提出了GeoSET,这是首个用于SET的全能模型,围绕一个单一的预训练父模型构建,该模型在通用协议下适应下游数据集。我们从超过1000万条SAR观测中精选了超过300万对高质量的SAR-EO图像对,涵盖不同的传感器、空间分辨率和地面采样距离。为了弥合SAR观测与预训练图像生成器之间的模态差距,我们开发了一个对斑点噪声鲁棒的SAR编码器,并在这一异构语料库上预训练条件生成器。所得的父模型通过低秩适应(LoRA)支持跨下游数据集的高效适应,仅更新生成器参数的0.60%,每个数据集大约需要一小时。在六个下游基准测试中,GeoSET在完全微调或LoRA下取得了FID和DISTS的最先进结果,展示了跨异构SAR-EO领域的有效迁移能力。
英文摘要
Paired synthetic aperture radar (SAR) and electro-optical (EO) imagery is increasingly available across sensors, resolutions, and geographic regions. Yet existing SAR-to-EO image translation (SET) methods are typically trained on a single, limited-scale dataset, producing models specialized to particular sensing conditions. We introduce GeoSET, the first generalist model for SET, built around a single pretrained parent that is adapted to downstream datasets under a common protocol. We curate over 3 million high-quality SAR--EO pairs from a collection of more than 10 million SAR observations, spanning diverse sensors, spatial resolutions, and ground sampling distances. To bridge the modality gap between SAR observations and a pretrained image generator, we develop a speckle-robust SAR encoder and pretrain the conditional generator on this heterogeneous corpus. The resulting parent supports efficient adaptation across downstream datasets through low-rank adaptation (LoRA), updating only 0.60% of the generator parameters and requiring approximately one hour per dataset. Across six downstream benchmarks, GeoSET achieves state-of-the-art results in FID and DISTS with full fine-tuning or LoRA, demonstrating effective transfer across heterogeneous SAR-EO domains.
CommentsPlease visit our project page https://kaist-viclab.github.io/GeoSET_site/