arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2501.12418cs.CVcs.AI

ImageRef-VL:在视觉语言模型中实现上下文图像引用

ImageRef-VL: Enabling Contextual Image Referencing in Vision-Language Models

  • University of Science and Technology of China(中国科学技术大学)
  • ByteDance(字节跳动)
  • Peking University(北京大学)
  • University of Oxford(牛津大学)
  • Tsinghua University(清华大学)

机构由 AI 辅助整理,请以论文原文为准。

Jingwei Yi, Junhao Yin, Ju Xu, Peng Bao, Yongliang Wang, Wei Fan, Hao Wang

更新

AI总结:

本文提出ImageRef-VL,通过指令微调增强开源视觉语言模型在对话中根据上下文引用相关图像的能力,实验显示性能提升88%。

AI中文摘要:

视觉语言模型(VLMs)在理解多模态输入方面展现了卓越的能力,并已被广泛集成到基于检索增强生成(RAG)的对话系统中。尽管当前由VLM驱动的聊天机器人能够在回答中提供文本来源引用,但它们在对话过程中引用上下文相关图像方面存在显著局限性。在本文中,我们引入了上下文图像引用——即根据对话上下文从检索文档中恰当地引用相关图像的能力——并系统地研究了VLM在这方面的能力。我们进行了首次针对上下文图像引用的评估,包括一个专门的测试数据集和评估指标。此外,我们提出了ImageRef-VL,一种通过在大规模、人工整理的多模态对话数据集上进行指令微调,显著增强开源VLM图像引用能力的方法。实验结果表明,ImageRef-VL不仅超越了专有模型,而且在上下文图像引用任务中,相较于最先进的开源VLM实现了88%的性能提升。我们的代码可在https://github.com/bytedance/ImageRef-VL获取。

英文摘要:

Vision-Language Models (VLMs) have demonstrated remarkable capabilities in understanding multimodal inputs and have been widely integrated into Retrieval-Augmented Generation (RAG) based conversational systems. While current VLM-powered chatbots can provide textual source references in their responses, they exhibit significant limitations in referencing contextually relevant images during conversations. In this paper, we introduce Contextual Image Reference -- the ability to appropriately reference relevant images from retrieval documents based on conversation context -- and systematically investigate VLMs' capability in this aspect. We conduct the first evaluation for contextual image referencing, comprising a dedicated testing dataset and evaluation metrics. Furthermore, we propose ImageRef-VL, a method that significantly enhances open-source VLMs' image referencing capabilities through instruction fine-tuning on a large-scale, manually curated multimodal conversation dataset. Experimental results demonstrate that ImageRef-VL not only outperforms proprietary models but also achieves an 88% performance improvement over state-of-the-art open-source VLMs in contextual image referencing tasks. Our code is available at https://github.com/bytedance/ImageRef-VL.

↑