arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.39128cs.CV

GeoGAT:双向时间采样结合层次图注意力用于全球视频地理定位

GeoGAT: Bidirectional Temporal Sampling Meets Hierarchical Graph Attention for Global Video Geo-localization

  • Information Engineering University(信息工程大学)
  • Zhengzhou University(郑州大学)

机构由 AI 辅助整理,请以论文原文为准。

Junchao Cui, Xuanzi Ma, Wenqi Shi, Hangyu Li, Biru Zhu, Chong Fu, Xiangyang Luo

AI总结:

GeoGAT通过双向时间采样与图注意力网络结合,并引入双重约束机制,解决了全球视频地理定位中多层级预测冲突问题,在CityGuessr68k和GeoGAT10k上均取得最先进性能。

AI中文摘要:

全球视频地理定位旨在推断视频在全球范围内的地理位置,并在四个地理层级(城市、州/省、国家和洲)上评估性能。现有方法通常采用单向均匀采样处理视频帧,并为每个层级训练独立的分类器,这导致关键地理线索的丢失以及层级间的预测冲突,尤其是在复杂的多镜头剪辑视频中。为解决这些局限,我们提出了GeoGAT,它将双向时间采样与图注意力网络(GATs)相结合。具体而言,GeoGAT提取正向和偏移反向的帧序列以构建互补的时空特征。这些融合特征随后被输入到预定义的地理层级图中,GATs在其中执行结构感知的消息传递,同时一个双重约束机制对预测进行剪枝以消除跨层级冲突。我们构建了GeoGAT10k数据集,包含来自全球166个城市的9,720个多镜头剪辑视频,专门用于基准测试在复杂视频结构上的泛化能力。在CityGuessr68k和GeoGAT10k上的实验结果表明,GeoGAT完全消除了层级冲突,并在所有四个地理层级上达到了最先进的性能。在CityGuessr68k上,GeoGAT在分类和检索两种协议下均优于最强基线,在城市层级上高出2.6个百分点。在更具挑战性的包含多镜头剪辑视频的GeoGAT10k上,准确率提升超过24个百分点,验证了对复杂现实场景的强泛化能力。

英文摘要:

Global video geo-localization aims to infer the geographic location of a video worldwide, evaluating performance across four geographic hierarchies: city, state/province, country, and continent. Existing methods typically employ one-way uniform sampling to process video frames and train independent classifiers for each hierarchy, which leads to the loss of key geographic cues and prediction conflicts between hierarchies, especially for complex multi-shot edited videos. To address these limitations, we propose GeoGAT, which integrates bidirectional temporal sampling with graph attention networks (GATs). Specifically, GeoGAT extracts forward and offset-reversed frame sequences to construct complementary spatiotemporal features. These fused features are then fed into a predefined geographical hierarchy graph, where GATs perform structure-aware message passing, while a dual-constraint mechanism prunes predictions to eliminate cross-hierarchy conflicts. We construct GeoGAT10k, comprising 9,720 multi-shot edited videos from 166 cities worldwide, specifically to benchmark generalization ability on complex video structures. Experimental results on CityGuessr68k and GeoGAT10k demonstrate that GeoGAT eliminates hierarchical conflicts entirely and achieves state-of-the-art performance across all four geographic hierarchies. On CityGuessr68k, GeoGAT outperforms the strongest baseline, evaluated under both classification and retrieval protocols, by 2.6 percentage points at the city level. On the more challenging GeoGAT10k with multi-shot edited videos, the accuracy improvement exceeds 24 percentage points, validating strong generalization to complex real-world scenarios.

↑