AI 中文总结
该研究针对遥感跨模态图像翻译的联合学习缺陷,提出解耦范式LTP-BIT,先学目标先验再做源域条件控制,以少量参数实现SOTA性能,且在少配对样本下保留高实例保真度。
AI 中文摘要
遥感跨模态图像翻译需保留源域观测内容,同时匹配目标域分布。现有方法从稀缺配对数据中联合学习目标先验与跨模态依赖,忽略了关键不对称性:仅后者本质需要跨模态对应。我们通过条件得分和去噪风险分析明确该区别,提出LTP-BIT(Learning the Target Priors Before Image Translation)这一先验优先范式,将两项学习任务解耦。LTP-BIT先从大规模非配对图像中学习目标域生成先验,再保留预训练骨干权重,通过P-DART(参数高效双流架构)学习源域条件控制。控制实验显示,先验匹配与缩放主要提升目标域真实感,实例保真度则更依赖条件适应。LTP-BIT在SAR转RGB、近红外转RGB基准上实现SOTA性能,仅用9.81%任务特定参数;在QXS-SAROPT数据集上,仅用25%配对样本即可保留近全数据的实例保真度。
英文摘要
Cross-modal image translation in remote sensing must preserve source-observed content while matching the target-domain distribution. Existing methods jointly learn the target prior and cross-modal dependence from scarce paired data, overlooking a key asymmetry: only the latter intrinsically requires cross-modal correspondence. We formalize this distinction through conditional-score and denoising-risk analyses and propose Learning the Target Priors Before Image Translation (LTP-BIT), a prior-first paradigm that decouples the two learning tasks. LTP-BIT first learns a target-domain generative prior from large-scale unpaired imagery, then retains the pretrained backbone weights and learns source-conditioned control through P-DART, a parameter-efficient dual-stream architecture. Controlled experiments show that prior matching and scaling primarily improve target-domain realism, whereas instance fidelity relies more strongly on conditional adaptation. LTP-BIT achieves state-of-the-art performance across SAR-to-RGB and NIR-to-RGB benchmarks using only 9.81% task-specific parameters. On QXS-SAROPT, it retains near-full-data instance fidelity with only 25% of the paired samples.
Comments26 pages, including supplementary material