arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SABRE:压力场景下视觉语言模型(VLM)的可扩展自动化基准测试

SABRE: Scalable and Automated Benchmarking of VLMs under Stress

Zixuan Lan, Luzhe Sun, Matthew R. Walter, Jiawei Zhou

arXiv 2608.07435首次发表:更新:

发表机构

University of Chicago; Toyota Technological Institute at Chicago; Stony Brook University(芝加哥大学; 芝加哥丰田技术学院; 石溪大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

SABRE是可扩展自动化VLM压力测试框架,可生成多维度测试样本,经实验验证其能有效评估VLMs对视觉证据与世界先验的依赖程度,支持多种压力测试场景。

AI 中文摘要

视觉语言模型(VLMs)发展迅速,但基准测试开发滞后,导致其弱点难以识别。构建压力测试成本高昂:样本需满足可控条件、保持可解答性并挑战当前模型。我们提出SABRE,一种可扩展的自动化流水线,可将测试 Primer(带有数据模式的Markdown任务设计)转换为结构化规范、生成或编辑后的图像以及问答对。自动过滤会移除被过滤VLM解决的候选样本,人工审核则验证候选样本有效性,支持标注修正和局部图像修复。我们实例化SABRE-Prior以测试VLMs是否遵循视觉证据而非依赖世界先验(对熟悉物体和场景的习得预期)。其600张图像和1000个问题涵盖上下文(熟悉场景中的意外实体)、纹理(反事实材料)、属性(非规范组件数量)及语言诱导(语言暗示但图像不支持的答案)。在6个VLMs上,宏平均准确率为17.8%至31.3%(均值22.6%)。真实图像属性控制对过滤VLM难度相当。SABRE-Counting和SABRE-Spatial试点显示该工作流支持其他压力测试设置。这些结果确立SABRE为构建和更新VLM压力测试的可复用框架,而非单一固定基准。

英文摘要

Vision-language models (VLMs) are improving rapidly, but benchmark development lags behind, making weaknesses hard to identify. Building stress tests is costly: samples must satisfy controlled conditions, remain answerable, and challenge current models. We present SABRE, a scalable, automated pipeline that converts a Test Primer (a Markdown Task Design with Data Schema) into structured specifications, generated or edited images, and question-answer pairs. Automated filtering removes candidates solved by a Filtering VLM, while human review verifies candidate validity and supports annotation correction and localized image repair. We instantiate SABRE-Prior to test whether VLMs follow visual evidence instead of relying on world priors -- learned expectations about familiar objects and scenes. Its 600 images and 1,000 questions span Context (unexpected entities in familiar scenes), Texture (counterfactual materials), Attribute (noncanonical component counts), and Language Elicitation (answers suggested by language but unsupported by the image). Across six VLMs, macro-average accuracy ranges from 17.8% to 31.3% (22.6% mean). A real-image Attribute control is comparably difficult for the Filtering VLM. SABRE-Counting and SABRE-Spatial pilots show that the workflow supports other stress-test settings. These results establish SABRE as a reusable framework for constructing and refreshing VLM stress tests rather than a single fixed benchmark.

Comments22 pages, 10 figures. Code and resources will be available at https://zesearch.github.io/vlm-SABRE/

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑