arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

GRASP:面向无人机视角细粒度跨模态理解的粒度感知区域对齐与语义原型学习

GRASP: Granularity-Aware Region Alignment and Semantic Prototype Learning for Fine-Grained Cross-Modal Understanding in Drone Views

Jiahui Cui, Yan Zhao, Kan Wei, Enze Zhu, Peirong Zhang, Lei Wang, Yiru Wang

arXiv 2608.09270首次发表:更新:

发表机构

Aerospace Information Research Institute, Chinese Academy of Sciences; University of Chinese Academy of Sciences; School of Electronic, Electrical and Communication Engineering, University of Chinese Academy of Sciences(中国科学院空天信息创新研究院; 中国科学院大学; 中国科学院大学电子电气与通信工程学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对无人机视角细粒度跨模态理解的背景杂波干扰与视觉同构歧义问题,提出GRASP框架,通过RFA和SPEM策略提升性能,在相关基准数据集上验证了有效性。

AI 中文摘要

无人机视角下的细粒度跨模态理解对空中视觉语言导航至关重要。然而,无人机场景固有的宽视场和俯视视角给视觉语言理解带来双重挑战:宏观层面,视觉表征中大量的背景杂波导致跨模态焦点错位,模型会优先考虑全局环境相似性而非特定对象细节;微观层面,视觉同构性造成歧义,候选对象共享相似几何结构却仅在细微属性上存在差异。为应对这些挑战,本文提出Granularity-Aware Region Alignment and Semantic Prototype(GRASP)学习框架,通过两种协同策略提升判别能力:具体而言,引入区域聚焦对齐(Region-Focused Alignment,RFA)以促进以对象为中心的跨模态对齐,同时抑制背景干扰;同时,为解决视觉同构问题,提出语义扰动增强匹配(Semantic Perturbation Enhanced Matching,SPEM),该方法利用经前景净化的语义原型码本(Semantic Prototype Codebook,SPC)构建语义扰动负样本,用于细粒度语义判别。在GeoText-1652基准数据集和未见过的ERA数据集上开展的大量实验表明,GRASP在无人机视角细粒度图像-文本检索任务中取得了具有竞争力的性能,验证了其对空中场景跨模态理解的有效性。本文的代码实现可在该https URL获取。

英文摘要

Fine-grained cross-modal understanding in drone views is essential for aerial vision-language navigation. However, the inherent wide field of view and overhead perspective of drone scenarios impose dual challenges on vision-language understanding. At the macro level, overwhelming background clutter in visual representations leads to Cross-Modal Focus Misalignment, where the model prioritizes global environmental similarities over specific object details. At the micro level, Visual Isomorphism creates ambiguity, where candidates share similar geometric structures yet differ only in subtle attributes. To address these challenges, we propose the Granularity-Aware Region Alignment and Semantic Prototype (GRASP) learning framework, enhancing discriminative capability through two synergistic strategies. Specifically, we introduce Region-Focused Alignment (RFA) to promote object-centric cross-modal alignment while suppressing background interference. Concurrently, to tackle visual isomorphism, we propose Semantic Perturbation Enhanced Matching (SPEM), which leverages a foreground-purified Semantic Prototype Codebook (SPC) to construct semantically perturbed negatives for fine-grained semantic discrimination. Extensive experiments on the GeoText-1652 benchmark and the unseen ERA dataset demonstrate that GRASP achieves competitive performance in drone-view fine-grained image-text retrieval, validating its effectiveness for cross-modal understanding in aerial scenarios. Our code implementation is available at https://github.com/UCAS-JC/GRASP.

CommentsAccepted at the 34th ACM International Conference on Multimedia (ACM Multimedia 2026, MM '26). 10 pages, 6 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑