PRISM-VLM:面向紧凑视觉语言模型的多轴判别基准
PRISM-VLM: A Multi-Axis Discriminative Benchmark for Compact Vision-Language Models
- NAVER Cloud AI(NAVER云AI)
- KAIST AI(韩国科学技术院AI)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对现有基准将紧凑视觉语言模型简化为单一准确率、难以区分模型差异的问题,提出多轴判别基准PRISM-VLM,沿七个轴评分并合成PScore,更可靠地区分模型并揭示行为差异,将发布完整流程与注释。
AI中文摘要:
紧凑型视觉语言模型(VLM)如今支撑着越来越多的多模态应用。然而,用于比较这些模型的基准却继承了面向前沿模型的设计理念:每个模型被简化为单一的准确率数值,这缩小了在饱和基准套件上的模型间差距,并将模型在更困难的基准上压入低分区间。我们提出PRISM-VLM,一个多轴判别基准,它沿七个轴对每个项目进行评分,覆盖了反复出现的失败模式(任务质量、行为鲁棒性和能力瓶颈),并将它们合并为一个单一的PScore,项目来自十五个公共基准。在过去的两年中,对于紧凑型VLM,在项目级配对自助法下,PScore比先前的单轴基准更可靠地区分模型对,并揭示了这些基准平均掉的行为差异。即使PScore在统计上无法区分的模型,在逐轴分布上也表现出显著差异,尤其是在谄媚(sycophancy)方面,这与单提示准确率几乎正交。我们将发布完整的流程、提示和逐项注释。
英文摘要:
Compact vision-language models (VLMs) now power a growing share of multimodal applications. The benchmarks used to compare them, however, inherit a frontier-centric design: each model is reduced to a single accuracy number, narrowing the inter-model gap on saturated suites and pressing models into low-score bands on harder ones. We introduce PRISM-VLM, a multi-axis discriminative benchmark that scores every item along seven axes covering the recurring failure modes (task quality, behavioral robustness, and capability bottlenecks) and combines them into a single PScore, with items recycled from fifteen public benchmarks. Across compact VLMs from the past two years, PScore separates model pairs more reliably than prior single-axis benchmarks under an item-level paired bootstrap, and surfaces behavioral differences these benchmarks average away. Even models with statistically indistinguishable PScores diverge sharply along the per-axis profile, particularly on sycophancy, which is nearly orthogonal to single-prompt accuracy. We will release the full pipeline, prompts, and per-item annotations.