AI 中文总结
研究针对SAR预训练中重建目标设计未确定的问题,提出应满足物理稳定性和语义尺度兼容性。通过SARATR-X-v2协调两者,其目标由固定结构提取器构建并融合成监督信号,在多个基准测试中性能最优,确立了预训练目标设计框架。
AI 中文摘要
掩蔽图像建模已成为合成孔径雷达(SAR)预训练的主导范式,但重建目标的设计仍未完全确定。本文认为,SAR预训练目标应满足两个条件以产生可转移的表示:一是基于物理的稳定性,即目标算子对相干成像中固有的乘性斑点具有近似不变性;二是语义尺度兼容性,即涵盖下游任务所需的异构空间尺度。这两个条件单独实现较容易,但联合起来则困难:基于物理的稳定性有利于固定算子,而语义尺度兼容性有利于数据驱动的合成。为此,SARATR-X-v2在单一设计中协调了这两者。目标通过跨越六个感受野的固定结构提取器构建,从盲点局部聚合到方向对数比区域对比度,并通过可学习权重融合成一个统一的监督信号用于掩蔽重建。在十二个分类、检测和分割的SAR基准上,SARATR-X-v2实现了最优的转移性能。在合成斑点变化下,相对于像素空间监督,所提出的目标将学习表示中的扰动漂移降低了近两个数量级。这些结果确立了基于物理的稳定性和语义尺度兼容性作为相干成像下预训练目标设计的原则框架,并表明有效的SAR预训练不是关于重建更多信号,而是关于重建正确的结构目标。
英文摘要
Masked image modeling has become a dominant paradigm for SAR pre-training, yet the design of the reconstruction target remains fundamentally unsettled. This article argues that a SAR pre-training target should satisfy two conditions to produce transferable representations: (i) physics-grounded stability, i.e., approximate invariance of the target operator to multiplicative speckle inherent in coherent imaging; and (ii) semantic scale compatibility, i.e., coverage of the heterogeneous spatial scales that downstream tasks demand. These two conditions are individually achievable but jointly difficult: physics-grounded stability favors fixed operators, while semantic scale compatibility favors data-driven composition. To this end, SARATR-X-v2 reconciles both within a single design. The target is constructed through fixed structural extractors spanning six receptive fields, from blind-spot local aggregation to directional log-ratio region contrast, and fused via learnable weights into one unified supervision signal for masked reconstruction. On twelve SAR benchmarks across classification, detection, and segmentation, SARATR-X-v2 achieves state-of-the-art transfer performance. Under synthetic speckle variation, the proposed target reduces perturbation drift in the learned supervision by nearly two orders of magnitude relative to pixel-space supervision. Taken together, these results support physics-grounded stability and semantic scale compatibility as a principled framework for pre-training target design under coherent imaging, and suggest that effective SAR pre-training is not about reconstructing more signal, but about reconstructing the right structural target.