基于序列观测的轨迹感知跨视角地理定位
Trajectory-aware Cross-view Geo-localization with Sequential Observations
- Washington University in St. Louis(圣路易斯华盛顿大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
研究跨视角地理定位,提出TrajLoc统一框架处理视频片段和路线描述,利用视觉与语言语义相互强化匹配,还提出TrajMod模块,实验显示该框架在视频和文本地理定位上比现有方法有显著优势。
AI中文摘要:
跨视角地理定位将地面观测与地理标记卫星图像进行匹配。近期方法表明视频片段等序列查询比单张图像能产生更丰富的时空线索,但忽略了路线描述这种互补的序列模态。为此引入SeqGeo - VL数据集及TrajLoc统一框架,它能处理视频片段和路线描述,通过利用视觉和语言语义相互强化跨视角匹配。还提出TrajMod模块。实验表明TrajLoc在视频和文本地理定位上比现有方法有显著提升。
英文摘要:
Cross-view geo-localization matches ground-level observations against geo-tagged satellite imagery. Recent methods show that sequential queries such as video clips yield richer spatiotemporal cues than single images, yet they overlook a complementary sequential modality: route descriptions -- which capture the same trajectory at a higher level of abstraction and are often the only input available (e.g., a user directing an autonomous vehicle to a pickup point). To bridge this gap, we introduce SeqGeo-VL, a dataset of $\sim$39K video-text-satellite triplets, and TrajLoc, a unified framework capable of processing both video clips and route descriptions. By leveraging both dense visual and abstract linguistic semantics, TrajLoc enables these modalities to mutually reinforce cross-view matching. We further propose TrajMod, a lightweight module that conditions query embeddings on trajectory geometry, yielding spatially-aware representations. Experiments show that TrajLoc achieves substantial gains over state-of-the-art methods on both video and text geo-localization. The project page is available at https://humblegamer.github.io/trajloc/.