基于结构化上下文推理提升基于知识的视觉问答性能
Boosting Knowledge-based Visual Question Answering with Structured Context Reasoning
查看机构详情
- School of Computer Science, Guangdong University of Technology(广东工业大学计算机学院)
- School of Artificial Intelligence, Shenzhen University(深圳大学人工智能学院)
- Graduate School of Advanced Science and Engineering, Hiroshima University(广岛大学先进科学与工程研究生院)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
本文提出SCoRe框架,通过上下文获取、选择、压缩三阶段处理多模态知识,在OK-VQA和A-OKVQA基准上性能优于现有最优方法,提升了基于知识的视觉问答效果。
中文摘要 AI 辅助
基于知识的视觉问答旨在通过将外部知识与视觉和文本信息相结合来回答关于图像的问题。近期方法常依赖上下文学习,以零样本或少样本方式用多模态上下文提示大语言模型(LLMs)。但我们发现,直接将异构视觉描述和检索到的知识拼接成长且非结构化的提示,会因存在过多无关上下文且缺乏显式关系结构而降低推理性能。本文提出一种基于LLM的结构化上下文推理(SCoRe)框架,该框架会为预测推断显式和隐式关系。SCoRe包含三个阶段:上下文获取阶段,生成多样视觉注释并通过高效两阶段多模态检索策略检索显式知识;上下文选择阶段,利用LLM引导的选择过滤相关的视觉、显式和隐式知识;上下文压缩阶段,执行关系逻辑蒸馏(RLD),将原始文本转换为显式实体-关系三元组,这些关系三元组可作为简洁且结构化的提示用于最终答案预测。在OK-VQA和A-OKVQA基准上开展的大量实验表明,SCoRe的性能始终优于现有最优方法。
英文摘要
Knowledge-based Visual Question Answering aims to answer questions about an image by integrating external knowledge with visual and textual information. Recent approaches often rely on in-context learning to prompt Large Language Models (LLMs) with multimodal context in a zero-shot or few-shot manner. However, we observe that directly concatenating heterogeneous visual descriptions and retrieved knowledge into long, unstructured prompts often degrades reasoning performance, due to both excessive irrelevant context and the lack of explicit relational structure. In this paper, we propose an LLM-based Structured Context Reasoning (SCoRe) framework that infers both explicit and implicit relationships for prediction. SCoRe consists of three stages: Context Acquisition, which generates diverse visual notes and retrieves explicit knowledge via an efficient two-stage multimodal retrieval strategy; Context Selection, which filters relevant visual, explicit, and implicit knowledge using LLM-guided selection; and Context Compression, which performs Relational Logic Distillation (RLD) to transform raw text into explicit entity-relation triplets. These relational triplets serve as a concise and structured prompt for final answer prediction. Extensive experiments on the OK-VQA and A-OKVQA benchmarks demonstrate that SCoRe consistently outperforms state-of-the-art methods.