arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

DiaVLo:诊断视觉-语言模型的行为

DiaVLo: Diagnosing Behaviours of Vision-Language Models

Lorenzo Corti, Jie Yang

arXiv 2609.22008首次发表:更新:

发表机构

Delft University of Technology(代尔夫特理工大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

DiaVLo是一个诊断框架,通过人工策展和生成能力构建行为规范,提供因果估计,识别VLM行为中的不一致,并揭示概念处理模式。

AI 中文摘要

视觉-语言模型(VLMs)依赖于在其子组件之间存储和传递适当的信息。验证VLM展现出期望的行为,同时避免有害行为,是其可靠部署的核心。然而,识别VLM行为的方法仍然稀缺。我们提出了DiaVLo,一个诊断框架,利用人工策展和VLM的生成能力来构建期望和观察到的VLM行为的规范,揭示潜在的不一致。除此之外,DiaVLo还提供因果估计,以识别引导VLM行为的最具影响力的概念。我们在多个开源VLM上,在分类和生成条件下评估了DiaVLo。我们的实验表明,DiaVLo产生的行为标签与模型性能相关,并为测量的性能提供背景。DiaVLo揭示了明确一致和不一致的行为,以及VLM感知、组织和优先处理概念的模式。

英文摘要

Vision-language models (VLMs) rely on storing and transferring appropriate information across their sub-components. Verifying that the VLMs exhibit desired behaviours, while avoiding harmful ones, is central to their reliable deployment. Yet, methods that identify VLM behaviours remain scarce. We present DiaVLo, a diagnostic framework that leverages human curation and VLMs' generation capabilities to construct specifications of desired and observed VLM behaviours, surfacing potential misalignments. Beyond this, DiaVLo also provides causal estimates to identify the most influential concepts steering VLM behaviours. We evaluate DiaVLo on several open-source VLMs under both classification and generation conditions. Our experiments show that DiaVLo produces behaviour labels that correlate with model performance and provide context for measured performance. DiaVLo surfaced behaviours that are clearly aligned and misaligned, alongside patterns in how VLMs perceive, organise, and prioritise concepts.

Comments34 pages. To appear in EMNLP 2026 (findings)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑