arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.18996cs.CVcs.AI

GrabVG:面向无人机图像视觉 grounding 的图注意力绑定方法

GrabVG: Graph-Attentive Binding for Visual Grounding in UAV Imagery

Chaowei Wang, Yan Di, Jingjun Sun, Baozhe Liu, Jiaxu Tian, Yuheng Li, Guangqian Guo, Shan Gao

首次发表
浏览论文内容

中文总结 AI 辅助

本文针对无人机图像视觉 grounding 的冗余与拓扑歧义问题,提出 GrabVG 框架,通过预注意假设搜索与图注意力特征绑定实现目标定位,在 AerialVG、AerialSense 数据集上精度显著优于基线且精度-速度权衡良好。

中文摘要 AI 辅助

无人机(UAV)图像中的视觉 grounding 旨在根据自然语言描述在复杂的鸟瞰视角场景中定位目标物体。然而,场景中大量小型、密集分布且视觉相似的物体造成了高视觉冗余,而重复的局部配置则引发了强烈的拓扑歧义。现有方法主要关注视觉-语言特征对齐或密集上下文交互,但难以区分实例间的细微差异,也无法有效利用空间拓扑结构,导致在高度拥挤场景中 grounding 结果不准确。为应对这些挑战,本文提出一种受人类视觉搜索启发的新型视觉 grounding 框架 GrabVG。GrabVG 将 grounding 明确分解为两个连续阶段:预注意假设搜索和图注意力特征绑定。具体而言,首先通过蒸馏引导的提议生成和文本感知的假设过滤生成一组紧凑的可靠物体假设,大幅减少背景干扰和语义不匹配;随后将这些假设组织成稀疏图,通过图注意力联合绑定并传播语言引导的实例内视觉线索与实例间拓扑关系,实现高效空间推理与准确目标定位。在 AerialVG 和 AerialSense 数据集上的大量实验表明,GrabVG 实现了良好的精度-速度权衡,Acc@0.5 分别达到 67.31% 和 80.34%,较对应基线分别提升了 10.55 和 8.76 个百分点。

英文摘要

Visual grounding in Unmanned Aerial Vehicle (UAV) imagery aims to localize a target object in complex bird's-eye-view scenes according to a natural language description. However, the abundance of small, densely distributed, and visually similar objects creates high visual redundancy, while repetitive local configurations give rise to strong topological ambiguity. Existing approaches mainly focus on visual--language feature alignment or dense contextual interaction, yet they struggle to distinguish subtle inter-instance differences and effectively exploit spatial topological structures, leading to inaccurate grounding in highly crowded scenarios. To address these challenges, we propose $\textbf{GrabVG}$, a novel visual grounding framework inspired by human visual search. GrabVG explicitly decomposes grounding into two sequential stages: $\textit{preattentive hypothesis search}$ and $\textit{graph-attentive feature binding}$. Specifically, we first generate a compact set of reliable object hypotheses through distillation-guided proposal induction and text-aware hypothesis filtering, substantially reducing background distractions and semantic mismatches. These hypotheses are then organized into a sparse graph, where language-guided intra-instance visual cues and inter-instance topological relationships are jointly bound and propagated via graph attention, enabling efficient spatial reasoning and accurate target localization. Extensive experiments on AerialVG and AerialSense show that GrabVG achieves a favorable accuracy--speed trade-off, reaching 67.31$\%$ and 80.34$\%$ Acc@0.5 and outperforming the corresponding baselines by 10.55 and 8.76 percentage points, respectively.

发表机构

  • Northwestern Polytechnical University(西北工业大学)
  • Harbin Institute of Technology(哈尔滨工业大学)
  • The Hong Kong Polytechnic University(香港理工大学)

机构由 AI 辅助整理,请以论文原文为准。

↑