arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.25422cs.AI

显著知识路径:用于高效知识密集型多模态问答的稀疏跨模态路由

Salient Knowledge Pathways: Sparse Cross-Modal Routing for Efficient Knowledge-Intensive Multimodal Question Answering

Noor Islam S. Mohammad, Uluğ Bayazıt

首次发表
浏览论文内容

中文总结 AI 辅助

研究知识密集型多模态问答,提出SKIP架构,通过问题引导视觉令牌修剪等方法,沿稀疏路径计算路由,结合自适应预算控制器,在五个基准测试中,以更少计算量和更低延迟达到或超越密集基线准确性。

中文摘要 AI 辅助

知识密集型多模态问答(KI-MMQA)涉及长视觉令牌序列、大型外部语料库上的密集检索和完全跨模态融合这三个昂贵的原语。现有系统对每个查询均统一承担这三项成本,而实际上只有一小部分视觉内容和检索到的知识与给定问题相关。我们引入了SKIP(显著知识注入路径),这是一种统一的推理架构,它根据问题、图像和难度估计,沿着稀疏路径进行计算路由。SKIP结合了问题引导的视觉令牌修剪、区域条件稀疏检索、二分稀疏交叉注意力和推测性知识验证,并配备了一个自适应预算控制器,根据预测的问题难度分配计算资源。我们推导出一个信息瓶颈界限,表明在现实的问题-图像互信息假设下,最优视觉稀疏率按$O(1/\sqrt{N})$缩放,并保证保留准确性。在五个KI-MMQA基准测试(OK-VQA、A-OKVQA、InfoSeek、Encyclopedic-VQA和ViQuAE)中,SKIP在使用的FLOP少$3.4$ - $6.8$倍且端到端延迟少$2.7$倍的情况下,匹配或超过了强大的密集基线的准确性。代码可在:此https URL获取。

英文摘要

Knowledge-intensive multimodal question answering (KI-MMQA) sits at the intersection of three expensive primitives: long visual token sequences, dense retrieval over large external corpora, and full cross-modal fusion. Existing systems pay all three costs uniformly per query, even though only a small fraction of visual content and retrieved knowledge is actually relevant to any given question. We introduce SKIP (Salient Knowledge-Injected Pathways), a unified inference architecture that routes computation along sparse pathways jointly conditioned on the question, the image, and a difficulty estimate. SKIP combines question-guided visual token pruning, region-conditional sparse retrieval, bipartite sparse cross-attention, and speculative knowledge verification with an adaptive budget controller that allocates compute proportional to predicted question difficulty. We derive an information-bottleneck bound showing that the optimal visual sparsity rate scales as $O(1/\sqrt{N})$ under realistic question-image mutual-information assumptions, with retained accuracy guarantees. Across five KI-MMQA benchmarks (OK-VQA, A-OKVQA, InfoSeek, Encyclopedic-VQA, and ViQuAE), SKIP matches or exceeds the accuracy of strong dense baselines while using $3.4$--$6.8\times$ fewer FLOPs and $2.7\times$ less end-to-end latency. Code available at: https://pmlrbd.github.io/skip/

发表机构

  • İTÜ(伊斯坦布尔技术大学)
  • Department of Computer Science, İTÜ(伊斯坦布尔技术大学计算机科学系)
  • Faculty of Computer Engineering, İTÜ(伊斯坦布尔技术大学计算机工程学院)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑