arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.04726cs.CV

桥接模态与任务:用于SAR到光学图像转换和语义分割的统一分层视觉Transformer

Bridging Modalities and Tasks: A Unified Hierarchical ViT for SAR-to-Optical Translation and Semantic Segmentation

Siyuan Liu, Xuze Zhang, Yongshun Wang, Licong Pan, Hang Liu, Huihui Li

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出BMT统一协同双任务学习框架,通过共享分层视觉Transformer联合优化SAR到光学图像转换与语义分割,在配对、非配对数据集上均取得良好效果,相关代码与数据集已公开。

中文摘要 AI 辅助

合成孔径雷达(SAR)图像具备全天候、昼夜观测能力,但相较于光学图像,其散斑噪声与非直观散射机制限制了图像的可解释性。用于SAR到光学(S2O)转换的生成模型可提升视觉可解释性,但现有方法常为追求视觉效果而忽略下游任务必需的语义结构约束。本文提出名为BMT(桥接模态与任务)的统一协同双任务学习框架,通过共享分层视觉Transformer联合优化S2O图像转换与语义分割任务。该框架包含四部分:(1)LocalViTBlock,通过可学习门控机制融合全局自注意力与空间深度卷积;(2)增强型输出模块,结合多尺度细化处理、色彩校正与抗锯齿技术,通过特征融合校准通道级色彩统计;(3)ControlNet式条件注入机制,将SAR小波特征与分割标签编码为多尺度特征金字塔,并通过零初始化卷积在每个编码器层注入;(4)有界Kendall不确定性加权方案,防止任一任务主导共享表示。我们在配对与非配对转换场景下分别开展评估,所用数据集为公开WHU-OPT-SAR配对数据集,以及由HRSID与DIOR构建的自建非配对舰船数据集。实验结果表明,所提方法兼具良好的S2O转换质量与语义分割性能,数据集与源代码已公开于指定链接。

英文摘要

Synthetic Aperture Radar (SAR) images have all-weather, day-and-night observation capabilities. However, compared with optical images, their speckle noise and non-intuitive scattering mechanism limit the interpretability of the images. Generative models for SAR-to-optical (S2O) conversion can improve visual interpretability, but existing methods often ignore the constraints on semantic structure, which are necessary for downstream tasks, for the sake of visual effects. We propose a unified collaborative dual-task learning framework, termed BMT (Bridging Modalities and Tasks), that jointly optimizes S2O image translation and semantic segmentation through a shared hierarchical Vision Transformer. The framework integrates: (1) a LocalViTBlock that fuses global self-attention with spatial depthwise convolution through a learnable gating mechanism; (2) an enhanced output module combining multi-scale refinement processing, color correction and anti-aliasing, which calibrates channel-level color statistics through feature fusion; (3) a ControlNet-style conditional injection mechanism that encodes SAR wavelet features and segmentation labels into a multi-scale feature pyramid and injects them at each encoder layer through zero-initialized convolution; (4) a bounded Kendall uncertainty weighting scheme that prevents either task from dominating the shared representation. We evaluate the framework under both paired and unpaired translation settings, on the public WHU-OPT-SAR paired dataset and a self-constructed unpaired ship dataset built from HRSID and DIOR, respectively. The experimental results show that the proposed method achieves competitive S2O translation quality and semantic segmentation performance. The dataset and source code have been publicly released at https://github.com/Lewisyuaner/BMT-S2O-main.

发表机构

  • School of Automation, Northwestern Polytechnical University(西北工业大学自动化学院)
  • School of Cybersecurity, Northwestern Polytechnical University(西北工业大学网络空间安全学院)

机构由 AI 辅助整理,请以论文原文为准。

↑