arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

用于多模态遥感图像生成的对比参数解缠

Contrastive Parameter Disentanglement for Multi-modal Remote Sensing Image Generation

Yu Zhang, Wenda Zhao, Haojun Tang, Haipeng Wang

arXiv 2607.23673首次发表:更新:

发表机构

School of Information and Communication Engineering, Dalian University of Technology; Unit 92728 of PLA(大连理工大学信息与通信工程学院; 中国人民解放军92728部队)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对现有遥感图像生成方法局限,提出对比参数解缠框架,通过对比参数解缠模块、解缠优化策略及查询-键结构转移机制,实现多模态遥感图像高质量生成,在相关任务中性能优越。

AI 中文摘要

现有遥感图像生成方法大多局限于单模态合成,无法利用多模态图像中的互补信息。为解决此局限,我们提出用于多模态遥感图像生成的对比参数解缠框架,能从单个文本提示生成跨光学、红外和合成孔径雷达(SAR)等多模态语义一致且结构对齐的图像。具体引入对比参数解缠模块在正交核心子空间参数层面解缠共享语义与模态特定属性,基于此开发解缠优化策略,先通过多模态对比目标约束LoRA适配器参数矩阵A捕获模态不变语义,再引导多个参数矩阵B在文本条件下学习模态特定属性。还设计查询-键结构转移机制确保生成图像结构对齐。实验表明该方法在生成质量、语义一致性和结构对齐方面优于现有方法,在下游目标分类任务中也性能优越。

英文摘要

Existing remote sensing image generation methods are largely confined to single-modality synthesis and therefore fail to exploit the complementary information inherent in multimodal imagery. To address this limitation, we propose a contrastive parameter disentanglement framework for multimodal remote sensing image generation, which generates semantically consistent and structurally aligned images across multiple modalities, including optical, infrared, and synthetic aperture radar (SAR), from a single text prompt. Specifically, we introduce a contrastive parameter disentanglement module that disentangles shared semantics from modality-specific attributes at the parameter level within an orthogonal core subspace. Based on this module, we develop a disentangled optimization strategy that first constrains the parameter matrix A of the LoRA adapter to capture modality-invariant semantics through a multimodal contrastive objective and then guides multiple parameter matrices B to learn modality-specific attributes under text conditioning. This strategy enables the simultaneous generation of multimodal images with consistent semantic content and distinct modality characteristics. Furthermore, to ensure structural alignment across the generated images, we devise a query-key structure transfer mechanism that jointly models multimodal sampling trajectories during inference by transferring structural correlation priors from an anchor modality to the remaining modalities. Extensive experiments demonstrate that our method outperforms state-of-the-art remote sensing image generation approaches in terms of generation quality, semantic consistency, and structural alignment, while also achieving superior performance in the downstream object classification task.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑