LoRetta:面向全球尺度遥感密集图像匹配的基础模型与大规模数据集
LoRetta: A Foundation Model and Extensive Dataset for Global-Scale Remote Sensing Dense Image Matching
浏览论文内容
中文总结 AI 辅助
针对全球尺度遥感密集图像匹配的挑战,本文提出结合可匹配性感知仿射定位与引导式密集配准的基础模型LoRetta,并构建含原生可匹配性标签的LEVIR-GM基准,实验表明其性能优于现有模型且迁移性良好。
中文摘要 AI 辅助
密集图像匹配建立像素级对应关系,是计算机视觉与摄影测量诸多应用的基础。然而将密集匹配扩展至全球尺度遥感领域仍面临挑战,因为图像对可能在采集时间、季节、视角、空间分辨率及地表覆盖状态上存在差异,由此产生的大幅几何偏移、部分重叠区域及本质上无法匹配的区域,使得直接预测密集对应关系既不可靠又低效。因此,我们将密集匹配重新表述为定位-配准问题:首先定位可匹配的重叠区域与仿射几何,再在对齐后的帧内优化密集残差。基于此表述,我们提出LoRetta,这是一个结合可匹配性感知仿射定位与引导式密集配准的基础模型;同时引入LEVIR-GM,这是一个全球尺度多时相光学匹配基准,包含数据集原生的可匹配性标签(10.3万对已对齐图像、82.7万对增强图像,覆盖六大洲、五年时间跨度,分辨率范围0.5-1024米)。我们还建立了统一的评估协议,用于稀疏、半密集及密集匹配器的评估。在LEVIR-GM上,LoRetta的曲线下面积(AUC)达到83.3%,优于最强基准模型RoMa v2 1.6个百分点;在1像素和2像素阈值下,其正确关键点百分比(PCK)分别提升6.5和8.2个百分点,同时推理延迟降低47.8%。宇航员对卫星及无人机(UAV)对卫星的地理定位实验进一步证明,LoRetta作为可复用几何对齐器具有良好的迁移性。
英文摘要
Dense image matching establishes pixel-wise correspondences and underpins broad applications in computer vision and photogrammetry. However, extending dense matching to global-scale remote sensing remains challenging because image pairs may differ in acquisition time, season, viewpoint, spatial resolution, and land-cover state. The resulting large geometric offsets, partial overlap, and intrinsically unmatchable regions make direct dense correspondence prediction unreliable and inefficient. We thus reformulate dense matching as localization-and-registration: first localizing the matchable overlap and affine geometry, then refining dense residuals within the aligned frame. Based on this formulation, we propose LoRetta, a foundation model coupling matchability-aware affine localization with guided dense registration. We also introduce LEVIR-GM, a global-scale multi-temporal optical matching benchmark with dataset-native matchability labels (103K aligned, 827K augmented pairs, six continents, five years, 0.5-1024 m resolution). We further establish a unified evaluation protocol for sparse, semi-dense, and dense matchers. On LEVIR-GM, LoRetta achieves an area under the curve (AUC) of 83.3%, outperforming the strongest baseline RoMa v2 by 1.6 points, with larger percentage of correct keypoints (PCK) gains of 6.5 and 8.2 points at 1 and 2 pixels, while reducing inference latency by 47.8%. Astronaut-to-satellite and unmanned aerial vehicle (UAV)-to-satellite geolocalization experiments further demonstrate its transferability as a reusable geometric aligner.
发表机构
- Beihang University(北京航空航天大学)
机构由 AI 辅助整理,请以论文原文为准。