面向自动驾驶的融合时空上下文建模的跨视角序列视觉定位
Cross-View Sequential Visual Localization with Spatio-Temporal Context Modeling for Autonomous Driving
AI总结:
本研究提出时序上下文增强的跨视角序列视觉定位框架,通过循环跨帧模块聚合历史上下文提升定位精度,在多个公开数据集及实装实验中均取得显著性能提升。
AI中文摘要:
连续可靠的定位是自动驾驶的核心需求。跨视角视觉定位将地面图像与卫星地图进行匹配,可为依赖全球导航卫星系统(GNSS)信号和高清(HD)地图的定位 pipeline 提供互补线索。现有多数跨视角视觉定位方法独立处理每一帧,未充分利用时序信息,在动态遮挡、光照变化和重复纹理场景下精度受限。本研究提出一种时序上下文增强的跨视角序列视觉定位框架,其中所提的循环跨帧模块聚合前序状态的历史上下文,以增强当前帧的粗粒度地面特征;这些增强特征用于卫星候选区域分类,而分层细粒度特征则支持精确的局部偏移估计。在 CVIS 数据集上,该方法将平均定位误差从 3.80 m 降至 1.57 m,且 R@1 m 指标从 8.14% 提升至 40.22%;直接迁移至 KITTI-CVL 数据集时,平均误差为 2.61 m,经目标域微调后平均误差进一步降至 2.27 m;在实装车辆的零样本野外实验中,平均误差达 2.84 m,R@5 m 达 96.86%。上述结果表明,时序上下文增强可显著提升跨视角定位精度,支持其在公开基准及实装道路上的稳健部署。
英文摘要:
Continuous and reliable localization is essential for autonomous driving. Cross-view visual localization matches ground images with satellite maps, providing complementary localization cues for pipelines that depend on Global Navigation Satellite System (GNSS) signals and high-definition (HD) maps. Most existing cross-view visual localization methods process each frame independently, leaving temporal information underused and limiting accuracy under dynamic occlusion, illumination variation, and repetitive textures. This study proposes a temporal-context-enhanced framework for cross-view sequence visual localization. The proposed recurrent cross-frame module aggregates historical context from the previous state to enhance the coarse ground feature of each current frame. These enhanced features facilitate satellite candidate-region classification, while hierarchical fine-grained features enable precise local offset estimation. On the CVIS dataset, the proposed method reduces mean localization error from 3.80 m to 1.57 m and increases R@1 m from 8.14% to 40.22%. Direct transfer to KITTI-CVL achieves a mean error of 2.61 m, with target-domain fine-tuning further reducing the mean error to 2.27 m. Zero-shot field experiments on a real-world vehicle achieve a mean error of 2.84 m and R@5 m of 96.86%. These results demonstrate that temporal context enhancement significantly improves cross-view localization accuracy and supports robust deployment on public benchmarks and real-world roads.