BRUCE:面向科学视觉-语言推理的污染升级下鲁棒性基准测试
BRUCE: Benchmarking Robustness Under Corruption Escalation for Scientific Vision-Language Reasoning
浏览论文内容
中文总结 AI 辅助
该研究提出BRUCE基准测试框架,通过RCI与T-RCI指标分析VLMs在科学视觉-语言推理任务中随污染升级的鲁棒性退化,实现可解释的失败分析。
中文摘要 AI 辅助
视觉-语言模型(VLMs)因输入图像质量较低或存在差异,在实际场景中常面临鲁棒性问题。本文旨在通过对输入图像施加模糊、低对比度等扰动与失真来分析VLMs的鲁棒性,为此提出BRUCE(污染升级下鲁棒性基准测试)——面向科学视觉-语言推理的多模态推理脆弱性框架。现有主流评估框架/研究主要关注干净任务的准确率,极少分析推理稳定性在鲁棒性维度的退化情况。除覆盖大范围输入扰动外,BRUCE采用两项新指标:鲁棒性污染指数(RCI)与遍历鲁棒性污染指数(T-RCI),用于量化多模态推理性能随视觉污染严重度在渐进式扰动缩放下的下降速率。我们在化学与数学推理任务的多个数据集上评估BRUCE,同时从四个高级推理领域分析污染诱导的预测失败:依赖OCR的推理、空间推理、符号推理及语义失败,每个领域包含细粒度的污染特定失败子类型,从而实现可解释的失败分析。
英文摘要
Visual-language models (VLMs) frequently struggle with robustness issues in real-world situations due to low- or varying-quality input images. In this paper, we aim at analyzing VLMs' robustness by applying perturbations and distortions to the input images, such as blur or low contrast. Toward this goal, we propose BRUCE (Benchmarking Robustness Under Corruption Escalation), a multimodal reasoning fragility framework for scientific vision-language reasoning. State-of-the-art evaluation frameworks/studies primarily focus on clean-task accuracy and rarely analyze how reasoning stability degrades across robustness dimensions. Besides varying over a wide-range of input perturbations, BRUCE employs two novel metrics -- Robustness Corruption Index (RCI) and Traversal-RCI (T-RCI) -- to quantify how rapidly multimodal reasoning performance deteriorates in VLMs as visual corruption severity increases under progressive perturbation scaling. We evaluate BRUCE across chemistry and mathematical reasoning tasks for multiple datasets, while analyzing corruption-induced prediction failures in terms of four high-level reasoning domains: OCR-dependent reasoning, spatial reasoning, symbolic reasoning, and semantic failures, with each containing fine-grained corruption specific failure subtypes, thereby enabling an interpretable failure analysis.
发表机构
- New York University Abu Dhabi(纽约大学阿布扎比分校)
机构由 AI 辅助整理,请以论文原文为准。