arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

UniHEAR:面向基于知识的视觉问答的统一异构源注意力检索

UniHEAR: Unified Heterogeneous-Source Attentive Retrieval for Knowledge-Based Visual Question Answering

Ganzhong Luo, Yang Ren, Hanyong Wang, Shuyu Zheng, Menglong Yang

arXiv 2608.01147首次发表:更新:

发表机构

School of Aeronautics and Astronautics, Sichuan University(四川大学航空航天学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对KB-VQA现有系统的单源检索瓶颈与检索源盲重排序问题,提出UniHEAR框架,在E-VQA和InfoSeek数据集上实现最优检索与VQA性能,提升Recall@1达6.7和1.2个点。

AI 中文摘要

基于知识的视觉问答(KB-VQA)需要从外部来源检索相关实体知识,以回答基于视觉的问题。现有检索增强系统存在两个关键局限:其一,依赖单一检索模态会造成单源检索瓶颈,遗漏仅能通过互补来源获取的真实实体;其二,双塔逐点重排序器存在检索源盲重排序问题,因它们忽略检索来源和候选级检索先验,导致冗余的模态依赖。为应对这些挑战,我们提出UniHEAR,一个用于异构源实体检索与重排序的统一轻量框架。UniHEAR为每个候选实体构建粗检索描述符,引入检索引导的注意力模态门控,基于该描述符调整模态注意力权重,辅以粗检索先验的熵加权源融合。结合对比学习与辅助模态保留损失的混合训练策略,在单个模型内统一实体级与部分级检索。在E-VQA和InfoSeek上的大量实验表明,UniHEAR实现了最优的检索和VQA性能,较最强基线分别提升了6.7和1.2个点的Recall@1,同时保持轻量的重排序架构。代码和模型可在this https URL获取。

英文摘要

Knowledge-Based Visual Question Answering (KB-VQA) requires retrieving entity knowledge from external sources to answer visually grounded questions. Existing retrieval-augmented systems suffer from two critical limitations. First, relying on a single retrieval modality creates a Single-Source Retrieval Bottleneck, missing ground-truth entities that are only accessible through complementary sources. Second, dual-tower pointwise rerankers suffer from Retrieval-Source-Blind Reranking, as they overlook retrieval origins and candidate-level retrieval priors, leading to redundant modality reliance. To address these challenges, we propose UniHEAR, a unified lightweight framework for heterogeneous-source entity retrieval and reranking. UniHEAR constructs a Coarse Retrieval Descriptor for each candidate entity, and introduces Retrieval-Guided Attentive Modality Gating to condition modality attention weights on this descriptor, complemented by Entropy-Weighted Source Fusion of coarse retrieval priors. A hybrid training strategy combining contrastive learning with an auxiliary modality-preserving loss unifies entity-level and section-level retrieval within a single model. Extensive experiments on E-VQA and InfoSeek demonstrate that UniHEAR achieves state-of-the-art retrieval and VQA performance, improving Recall@1 by 6.7 and 1.2 points over the strongest baselines while maintaining a lightweight reranking architecture. Code and model are available at https://github.com/iven-luo/UniHEAR.

CommentsAccepted by ACM MM 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑