发表机构
University of Idaho; University of California, Santa Barbara; Appalachian State University(爱达荷大学; 加州大学圣巴巴拉分校; 阿巴拉契亚州立大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出无需训练、无需GPU的SCDF自校准密集位移场,通过预测-测量-过滤循环实现大型光学卫星影像可靠配准,在584组多源影像对测试中零失败,显著降低了端点误差。
AI 中文摘要
配准是光学卫星影像几乎所有多时相和多传感器应用的基础,但业务化产品仍存在远高于亚像素级的记录偏移,而亚像素级偏移会导致变化检测、时间序列分析和数据融合性能下降。真实影像对会在多个维度上存在差异,包括传感器响应、场景内容、观测几何、分辨率、拼接缝等,且拼接缝并非单一全局运动。现有工具嵌入了针对其开发数据调整的运动模型和常数,适配的影像对可实现精确配准,不适配的则要么匹配失败,要么返回数十像素的错误结果且未报告失败。学习型匹配器需依赖GPU,且在训练分布之外无法保证精度。本文提出SCDF(自校准位移场,self-calibrating displacement fields),这是一种无需训练、无需GPU的估计器,其运动模型为逐像素密集位移场本身,因此不存在场景运动超出模型范围的情况。该方法在分辨率金字塔上运行单一的预测-测量-过滤循环:累积位移场预测待配准影像的每个补丁在参考影像中的位置,通过RootSIFT匹配和相关计算将位移测量至亚像素精度,所有阈值均针对影像对本身校准的滤波器决定保留哪些位移。一种无需针对每个数据集调整的配置可在单个CPU核心上处理完整的8192²场景。在由真实Sentinel-2、Landsat-8/9和NAIP影像构建的584组地面真值影像对上,对比7种经典基线和2种零样本预训练匹配器,SCDF实现了所有影像对的零失败配准,将最佳基线的真实影像对中位数端点误差从6.83米降至4.17米,将其90百分位误差从17.8米降至7.77米。
英文摘要
Co-registration underlies nearly every multi-temporal and multi-sensor use of optical satellite imagery, and operational products still carry documented offsets well above the fraction-of-a-pixel scale at which change detection, time series, and data fusion degrade. Real image pairs differ along several axes at once (sensor response, scene content, viewing geometry, resolution, mosaic seams), and the last of these is not a single global motion. Existing tools embed a motion model and constants tuned to their development data; a pair that fits is registered precisely, while one that does not either fails to match or returns a result wrong by tens of pixels with no failure reported. Learned matchers add a GPU requirement and carry no accuracy guarantee outside their training distribution. We present SCDF (self-calibrating displacement fields), a training-free, GPU-free estimator whose motion model is the dense per-pixel displacement field itself, so no scene motion falls outside the model. A single predict--measure--filter loop runs over a resolution pyramid: the accumulated field predicts where each patch of the moving image falls in the reference, RootSIFT matching and a correlation pass measure the displacement there to sub-pixel precision, and filters whose thresholds are all calibrated on the image pair itself decide what survives. One configuration, with no per-dataset tuning, processes full $8192^2$ scenes on a single CPU core. On 584 constructed-ground-truth pairs built from real Sentinel-2, Landsat-8/9, and NAIP imagery, against seven classical baselines and two zero-shot pretrained matchers, SCDF registers every pair with zero failures, reduces the best baseline's real-pair median end-point error from 6.83 to 4.17m, and cuts its 90th percentile from 17.8 to 7.77m.