发表机构
China University of Mining and Technology; University of Ottawa(中国矿业大学; 渥太华大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对遥感图像中有向目标视觉定位,提出O$^2$-VG模型家族(含Transformer、通用提案预测和视觉语言模型)及DIOR-R-RSVG数据集,在多个基准上表现优越。
AI 中文摘要
遥感图像中的视觉定位旨在定位由指代表达式描述的目标。大多数现有方法预测水平边界框,这对于任意方向的目标往往不准确。为解决这一局限,我们引入了O$^2$-VG,一个用于有向目标视觉定位的模型家族,包含三种互补设计。具体而言,O$^2$-VG-Trans是一种用于有向目标视觉定位的跨模态Transformer,它为模型家族建立了强大的判别基础。在此基础上,O$^2$-VG-Uni预测可能的背景目标的通用有向提案,无需特定文本提示,它还通过缓存的提案嵌入支持目标检索。使用这些通用有向提案作为输入提示,O$^2$-VG-VLM是一种自回归视觉语言模型,通过多标记预测并行生成有向框标记块。此外,我们构建了DIOR-R-RSVG,一个用于遥感图像中有向目标视觉定位的数据集,它提供图像、表达式和有向框三元组用于训练和评估。总之,O$^2$-VG家族提供了一个灵活框架,涵盖判别性Transformer和生成式视觉语言模型,在多个基准上取得了优越性能。代码可在https://github.com/wokaikaixinxin/ai4rs获取。
英文摘要
Visual grounding in remote sensing images aims to locate objects described by referring expressions. Most existing methods predict horizontal bounding boxes, which are often inaccurate for objects with arbitrary orientations. To address this limitation, we introduce O$^2$-VG, a family of models for oriented object visual grounding with three complementary designs. Specifically, O$^2$-VG-Trans is a cross-modality transformer for oriented object visual grounding. It establishes a strong discriminative foundation for the model family. Building upon it, O$^2$-VG-Uni predicts universal oriented proposals for possible foreground objects without specific text prompts. It also supports object retrieval through cached proposal embeddings. Using these universal oriented proposals as input prompts, O$^2$-VG-VLM is an autoregressive vision-language model. It generates oriented box token blocks in parallel through multi-token prediction. In addition, we construct DIOR-R-RSVG, a dataset for oriented object visual grounding in remote sensing images. It provides image, expression, and oriented box triplets for training and evaluation. Together, the O$^2$-VG family provides a flexible framework that spans discriminative transformers and generative vision-language models. It achieves superior performance across multiple benchmarks. Code is available at https://github.com/wokaikaixinxin/ai4rs and https://github.com/wokaikaixinxin/Eagle_o2_vg.