arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SAREO-FM:用于SAR-EO基础模型的解耦语义监督

SAREO-FM: Decoupled Semantic Supervision for SAR-EO Foundation Models

Jeonghyeok Do, Munchurl Kim

arXiv 2610.09317首次发表:更新:

发表机构

Korea Advanced Institute of Science and Technology (KAIST)(韩国科学技术院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

SAREO-FM通过解耦语义与重建监督,利用模态令牌和语义查询,在SAR-1M上预训练,实现SAR和EO单模态及联合感知的强迁移性能。

AI 中文摘要

合成孔径雷达(SAR)和光电(EO)图像提供互补的观测:SAR能够实现昼夜、抗天气的感知,而EO则提供丰富的外观和细粒度的语义线索。我们引入了SAREO-FM,它避免了强制单一令牌流承担两个不同角色:模态令牌通过掩码重建保留每个传感器如何观测场景,而可学习的语义查询在预训练视觉基础模型(VFM)的引导下捕获场景包含的内容。通过将查询与SAR和EO令牌联合编码,查询获得基于模态的语义上下文,而模态令牌输出仍然是掩码重建的显式目标。这种设计将语义和重建监督分配给不同的令牌流,同时保持它们在共享编码器内的交互。在百万级SAR-1M语料库上预训练,SAREO-FM对仅SAR和仅EO输入都实现了强大的单模态迁移,同时在受益于互补感知的任务上,通过联合SAR-EO观测带来了显著的增益。

英文摘要

Synthetic aperture radar (SAR) and electro-optical (EO) imagery provide complementary observations: SAR enables day-and-night, weather-resilient sensing, whereas EO provides rich appearance and fine-grained semantic cues. We introduce SAREO-FM, which avoids forcing a single token stream to serve two distinct roles: modality tokens preserve how each sensor observes the scene through masked reconstruction, while learnable semantic queries capture what the scene contains under guidance from a pretrained vision foundation model (VFM). By jointly encoding these queries with SAR and EO tokens, the queries acquire modality-grounded semantic context, while the modality-token outputs remain the explicit targets of masked reconstruction. This design assigns semantic and reconstruction supervision to separate token streams while preserving their interaction within the shared encoder. Pretrained on the million-scale SAR-1M corpus, SAREO-FM achieves strong unimodal transfer for both SAR-only and EO-only inputs, while delivering substantial gains from joint SAR--EO observations on tasks that benefit from complementary sensing.

CommentsPlease visit our project page at https://kaist-viclab.github.io/SAREO-FM_site/

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑