arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.33694cs.CV

看见与求解不足以应对视觉语言模型

Seeing and Solving Are Not Enough for Vision-Language Models

  • Sun Yat-sen University(中山大学)
  • Zhejiang University(浙江大学)
  • The Hong Kong University of Science and Technology(香港科技大学)
  • Institute of Information Engineering(信息工程研究所)
  • Hunan University(湖南大学)

机构由 AI 辅助整理,请以论文原文为准。

Ziheng Wang, Mingxuan Xie, Yilin Liu, Dayan Wu, Yang Li, Pengwen Dai

AI总结:

本文发现视觉语言模型即使分别具备视觉提取和求解能力,仍可能组合失败,提出状态实现微调(SRT)方法,通过先输出任务状态再回答,显著提升多模态回答准确率。

AI中文摘要:

视觉语言模型(VLMs)通过结合视觉信息提取与下游问题求解来回答视觉问题。我们研究一个基本问题:错误的答案是否必然反映视觉提取或问题求解的失败?一个模型可能在分别测试时两种能力都成功,但在原始多模态问题上仍然失败,这种区别是整体答案准确率无法揭示的。为了研究这一点,我们在多个VLM和视觉领域进行了问题级别的实证分析。我们定义了一个可精确评分的任务状态(即足以解答问题的视觉信息),并用它来测试同一模型是否能提取所需状态、从真实状态求解问题以及回答原始多模态问题。我们发现,组合失败(即提取和求解都成功但直接回答失败)在多个VLM和数据集上占直接回答错误的17.7%至75.6%。为了解决这种失败模式,我们引入了一种简单而有效的方法,称为状态实现微调(SRT)。SRT在保持预训练VLM权重冻结的同时,微调附加到语言模型层的LoRA适配器。它训练模型在单个自回归响应中输出真实任务状态后再给出最终答案。SRT相比标准监督微调提高了1.7至14.1个百分点,并修复了92.5%至98.1%的诊断出的组合失败。使用SRT训练的单个LoRA适配器还在显著不同的任务状态结构上提升了性能。我们的工作表明,同时具备视觉提取和问题求解能力并不能保证正确的多模态回答。要求模型先输出解答问题所需的视觉信息有助于弥合这一差距。

英文摘要:

Vision-language models (VLMs) answer visual questions by combining visual information extraction with downstream problem solving. We investigate a fundamental question: Does an incorrect answer necessarily reflect a failure in visual extraction or problem solving? A model may succeed at both abilities when tested separately yet still fail on the original multimodal question, a distinction that overall answer accuracy cannot reveal. To study this, we perform a question-level empirical analysis across multiple VLMs and visual domains. We define an exactly scorable task state (i.e., the visual information sufficient to solve a question) and use it to test whether the same model can extract the required state, solve the question from the ground-truth state, and answer the original multimodal question. We find that composition failures, where extraction and solving both succeed but direct answering fails, account for 17.7% to 75.6% of direct-answering errors across multiple VLMs and datasets. To address this failure mode, we introduce a simple yet effective method, termed State Realization Tuning (SRT). SRT fine-tunes LoRA adapters attached to the language-model layers while keeping the pretrained VLM weights frozen. It trains the model to output the ground-truth task state before the final answer in a single autoregressive response. SRT improves over standard supervised fine-tuning by 1.7 to 14.1 percentage points and repairs 92.5% to 98.1% of diagnosed composition failures. A single LoRA adapter trained with SRT also improves performance across substantially different task-state structures. Our work shows that having both visual extraction and problem-solving capabilities does not guarantee correct multimodal answering. Requiring the model to first output the visual information needed to solve the question can help bridge this gap.

补充信息

↑