arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.33855cs.CVcs.AIcs.CLcs.LGcs.NE

程序验证的视觉语言模型自我进化

Program-Verified Self-Evolution for Vision-Language Models

Ahmed Heakl, Sungik Choi, Moontae Lee, Salman Khan

首次发表
浏览论文内容

中文总结 AI 辅助

针对视觉语言模型自我进化中标签错误率高的问题,提出VQS方法,通过程序验证结构化记录生成问题与答案,显著提升准确率,并在多个基准上超越基线。

中文摘要 AI 辅助

自我进化的视觉语言模型在未标注图像上训练由它们自己生成的问题。由于这些问题没有标准答案,先前的方法通过多数投票或模型评判器对采样答案进行标注。在人工评估中,我们发现自我进化过程中产生的多数投票标签中有24%是错误的,模型评判器标签中有18%是错误的。为了解决这个问题,我们提出了用于自我进化模型的可验证问答生成方法(VQS),它改变了模型评判答案的方式。模型不再对答案进行投票,而是将每张图像解析为结构化记录,例如场景图、图表表格或图表图。固定程序随后根据记录编写问题并计算答案。模型仍然充当视觉检查器,但只逐条确认程序读取的单个事实,每次一个简短声明。这些声明级别的检查选择解析器的训练目标,因此解析器也在没有标签的情况下得到改进。人工评估者发现VQS答案的正确率为94%,而多数投票为76%。在十个基准测试中,VQS在2B、4B和8B规模下将Qwen3-VL提升了最多3.18分,并且在每个规模上都优于最强的自我进化基线。在三轮训练中,增益持续增长,在2B规模达到3.84分。代码已在此https URL发布。

英文摘要

Self-evolving vision-language models train on questions they generate from unlabeled images. Since these questions have no gold answers, prior methods label them by majority vote over sampled answers or by a model judge. In a human evaluation, we find that 24\% of majority-vote labels and 18\% of model-judge labels produced during self-evolution are wrong. To address this problem, we present Verifiable QA Generation for Self-Evolving Models (VQS), which changes how the model judges answers. Instead of voting on an answer, the model parses each image into a structured record, such as a scene graph, a chart table, or a diagram graph. Fixed programs then write a question from the record and compute its answer. The model still acts as a visual checker, but it only confirms the individual facts the program reads, one short claim at a time. These claim-level checks select the parser's training targets, so the parser also improves without labels. Human raters find 94\% of VQS answers correct, against 76\% for majority voting. Across ten benchmarks, VQS improves Qwen3-VL by up to 3.18 points at the 2B, 4B, and 8B scales and outperforms the strongest self-evolving baseline at each. Gains keep growing over three training rounds, reaching 3.84 points at 2B. Code is released at https://github.com/ahmedheakl/VQS

发表机构

  • LG AI Research(LG AI研究院)
  • MBZUAI(穆罕默德·本·扎耶德人工智能大学)
  • Australian National University(澳大利亚国立大学)
  • University of Illinois at Chicago(伊利诺伊大学芝加哥分校)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑