arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

VisionQ:用于计算机视觉定性分析的VLM-as-a-Judge分类体系、数据集与基准

VisionQ: VLM-as-a-Judge Taxonomy, Dataset and Benchmark for Qualitative Analysis in Computer Vision

Vu Dinh Xuan, Duc-Hai Nguyen, Minh-Dung Dao, Vu Quynh Giao, Quang Hong Nguyen, Binh-Son Hua, Barry O'Sullivan, David Murphy, Hoang D. Nguyen

arXiv 2610.00666首次发表:更新:

发表机构

University of Information Technology, VNU-HCM; University College Cork; Hanoi University of Science and Technology; Trinity College Dublin(越南国立胡志明市大学信息技术大学; 科克大学; 河内科技大学; 都柏林圣三一大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

VisionQ构建了首个基于同行评审比较图的VLM评判基准,通过标准条件视觉判别任务,提出分类体系、数据集和DPO调优的评判模型,显著提升评判准确率并揭示现有模型的不足。

AI 中文摘要

定性比较图是计算机视觉论文中的核心证据,而视觉语言模型(VLMs)正越来越多地被用于评判这些图。然而,现有基准仅对标量质量或整体偏好进行评分,因此评判者可能因选择了错误的视觉理由而选中了偏好图像,却仍获得奖励。我们引入了VisionQ,这是首个基于同行评审的CV比较图构建的基准,它使每个判断都基于一个明确的视觉标准:每个问题都陈述了该标准,只有当评判者选择了作者在该标准上认定为最佳的输出时,它才会获得认可。我们将此任务称为“标准条件视觉判别”。VisionQ包含:(1)一个包含1,409篇CVPR和ICCV论文的语料库,其中包含1,800多个经过验证的比较图和3,911个手工标注的数据点,将方法裁剪与作者陈述的视觉声明联系起来;(2)一个六轴、51叶的视觉标准分类体系,用于定性判断背后的视觉标准;(3)一个标准条件评估协议,隐藏方法名称、标题和论文身份,并报告每个标准的准确率;以及(4)VisionQ-Judge,一个基于对称证据对训练的DPO调优的Gemma-4-E4B评判模型,在保留测试集上将最后选项预测减少了7.0个百分点,并将准确率提高了2.5个百分点。通过评估20个开源和闭源VLM评判模型,我们发现最强的模型仅达到63.1%的准确率(随机水平为32.2%),且可靠性在不同标准间差异显著。代码:此https URL。数据:此https URL。

英文摘要

Qualitative comparison figures are central evidence in computer vision papers, and vision-language models (VLMs) are increasingly used to judge them. Yet existing benchmarks score only scalar quality or overall preference, so a judge can be rewarded for picking the preferred image for the wrong visual reason. We introduce VisionQ, the first benchmark built from peer-reviewed CV comparison figures that grounds every judgment in a named visual criterion: each question states the criterion, and a judge is credited only when it selects the output the authors identify as best on that criterion. We call this task criterion-conditioned visual discrimination. VisionQ comprises (1) a corpus of 1,409 CVPR and ICCV papers with 1,800+ validated comparison figures and 3,911 hand-annotated data points linking method crops to author-stated visual claims; (2) a six-axis, 51-leaf taxonomy of the visual criteria behind qualitative judgment; (3) a criterion-conditioned evaluation protocol that hides method names, captions, and paper identity and reports accuracy per criterion; and (4) VisionQ-Judge, a DPO-tuned Gemma-4-E4B judge trained on symmetric evidence pairs, which reduces last-option predictions by 7.0pp and improves accuracy by 2.5pp on a held-out test set. Evaluating 20 open- and closed-source VLM judges, we find that the strongest reach only 63.1% accuracy (chance 32.2%) and that reliability varies sharply across criteria. Code: https://github.com/ReML-AI/visionq. Data: https://huggingface.co/datasets/visionq-anon-2026/VisionQ-1k.

Comments29 pages, 18 figures, 6 tables. Code: https://github.com/ReML-AI/visionq

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑