发表机构
The University of Tokyo; Nara Institute of Science and Technology (NAIST)(东京大学; 奈良科学技术研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出结合检索增强生成(RAG)的框架,构建艺术品知识图谱ExpArt-KG,通过交替生成答案与检索知识、用正确性判断控制搜索,提升艺术品图像解释的详细度并降低检索成本。
AI 中文摘要
大型视觉语言模型(LVLMs)在图像文本生成和视觉问答任务中表现出色,但难以全面准确描述图像中物体相关实体与概念间的事实关系。本研究提出一种利用检索增强生成(RAG)高效挖掘知识图谱事实信息的框架,旨在让LVLMs生成详细准确的图像解释。具体而言,该方法在答案生成与知识图谱检索间交替进行,通过正确性判断控制搜索,高效获取必要且充足的事实信息。研究还构建了艺术品领域的知识图谱ExpArt-KG,其中图像与实体的对应关系明确。将所提方法应用于该知识图谱,实验表明其提升了艺术品解释的详细程度,在保持与固定次数迭代相当的生成质量的同时,降低了外部知识的检索成本。
英文摘要
Large Vision-Language Models (LVLMs) achieve strong performance on image-grounded text generation and visual question answering. However, it remains difficult for them to comprehensively and accurately describe the factual relations among the entities and concepts associated with the objects depicted in an image. In this work, we propose a framework that efficiently exploits factual information from a knowledge graph via retrieval-augmented generation (RAG), with the goal of enabling LVLMs to generate detailed and accurate image explanations. Specifically, our method alternates between answer generation and knowledge-graph retrieval, and controls the search using a correctness judgment, thereby acquiring the necessary and sufficient factual information efficiently. We also construct a knowledge graph for the artwork domain (ExpArt-KG), in which the correspondence between images and entities is unambiguous. Applying the proposed method to this knowledge graph, we show experimentally that it improves the level of detail of artwork explanations and reduces the retrieval cost of external knowledge while maintaining generation quality comparable to that of iterating a fixed number of times.