arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.11126cs.CVcs.AIcs.ET

超越基准:利用视觉语言模型揭示真实世界条件下的系统性分类失败

Beyond Benchmarks: Using VLMs to Reveal Systematic Classification Failures Under Real World Conditions

  • TNO(荷兰应用科学研究组织)

机构由 AI 辅助整理,请以论文原文为准。

Dieuwertje Alblas, Alma M. Liezenga, Jan Erik van Woerden, Fedor Taggenbrock, Dalia Aljawaheri, Klamer Schutte

AI总结:

本研究提出基于视觉语言模型的错误切片检测方法,用于加速分类模型验证与确认,初步评估其在国防领域的适用性,虽未达全自动化,但显示出加速V&V过程的潜力。

AI中文摘要:

分类模型的验证与确认(V&V)对于实现广泛的传感器处理应用至关重要。目前,V&V过程依赖于耗时的人工检查错误样本来寻找有意义的模式。本研究探索使用视觉语言模型(VLMs)来加速这一繁琐过程。VLMs经过训练,将图像嵌入到语义上有意义的向量表示中,从中可以提炼出人类可解释的系统性错误。在国防背景下部署这种基于VLM的方法引入了两大挑战:(1)国防领域在VLMs的训练数据中代表性不足,(2)周围环境和背景的多样性不如其他领域。本研究对基于VLM的方法在国防应用V&V中的适用性进行了初步评估。我们提出了一种基于VLM的错误切片检测(ESD)方法,该方法独立地对分类模型产生的系统性错误进行分组和标记。我们证明,该方法能够识别非军事数据集中操作相关的、人为添加的扰动。在军事背景下,我们的方法根据周围环境对图像进行聚类和描述,但聚类描述之间也存在重叠。我们进一步研究了军事与非军事数据集之间嵌入变异的差异,这仍是一个值得关注的话题。尽管结果尚不足以支持通过基于VLM的ESD实现完全自动化的V&V,但它们表明VLMs未来可用于加速V&V过程。

英文摘要:

Verification and validation (V&V) of classification models is crucial to enable a wide range of sensor processing applications. Currently, the V&V process relies on time-consuming manual inspection of erroneous samples to find meaningful patterns. This work explores the use of Vision Language Models (VLMs) to speed up this laborious process. VLMs are trained to embed images into a semantically meaningful vector representation, from which human-interpretable systematic errors can be distilled. Deploying such VLM-based methods in a defence context introduces two major challenges: (1) the defence domain is underrepresented in the training data of VLMs, and (2) surroundings and context are less diverse than for other domains. This study provides an initial assessment of the suitability of VLM-based methods for V&V of defence applications. We propose a VLM-based error slice detection (ESD) method that independently groups and labels systematic errors made by a classification model. We demonstrate that this method is able to identify operationally-relevant artificially added perturbations in a non-military dataset. In a military context, our method clusters and describes images based on their surroundings, but also exhibits overlap between cluster descriptions. We further investigate the difference in embedding variation between our military and non-military dataset, which remains a topic of interest. Although the results do not yet warrant fully automated V&V through VLM-based ESD, they show that VLMs could be used to accelerate V&V processes in the future.

补充信息

↑