发表机构
University of Gothenburg; Chalmers University of Technology(哥德堡大学; 查尔姆斯理工大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文针对汽车感知系统中视觉语言模型的幻觉问题,提出一种结合本体标注与同义词评估的设计时资质评估工作流,以补充运行时监控,实现对三种VLM幻觉的确定性量化,支持模型比较与部署决策。
AI 中文摘要
人工智能领域已被广泛应用于众多应用领域。视觉语言模型(Vision Language Models, VLMs)是近期发展起来的先进人工智能技术之一,已被探索用于支持车辆感知、安全保证等汽车功能。然而,此类语言模型容易产生幻觉(hallucinations),对可能集成它们的汽车系统的安全性构成潜在威胁。在汽车领域,VLM不仅可能幻觉出交通物体,还可能无法识别实际存在的交通物体,这可能导致危险情况的发生。尽管我们观察到越来越多的文献提出用于安全可信人工智能的验证与确认(verification and validation, V&V)技术,但这些方法往往被孤立地研究,要么侧重于运行时(run-time)阶段,要么侧重于设计时(design-time)阶段。在诸如汽车感知系统等安全关键、现实场景中,这种孤立的技术可能是不充分的。在本文中,我们基于Huang等人提出的分类法,分析了设计时和运行时的验证与确认技术。我们提出了一项汽车研究,其中提出了一种设计时资质评估工作流(qualification workflow),以补充运行时监控。该工作流结合了基于固定安全相关本体(ontology)的结构化标注系统,以及基于同义词的评估过程,以统计方式评估三种最先进的VLM在nuScenes数据集上的表现。我们观察到,所提出的技术能够对VLM在汽车感知相关任务中产生的幻觉进行确定性和可重复的量化。所提出的工作流支持设计时验证与确认过程中的模型比较和面向部署的工程决策,并将有助于形成一种整体性的验证策略,以努力实现可信赖的汽车感知系统。
英文摘要
The field of Artificial Intelligence has been adopted for many application domains. Vision Language Models are one of the recently advanced AI techniques that have been explored to support automotive features such as vehicle perception, and safety assurance. However, such language models are prone to hallucinations, posing a potential threat to the safety of automotive systems that may incorporate them. Within the automotive domain, VLMs could not only hallucinate traffic objects, but could also fail to identify traffic objects that are actually present, which may potentially lead to dangerous situations. Though we have observed a growing body of literature that proposes verification and validation techniques for safe and trustworthy AI, these methods are often studied in isolation, focusing either on run-time or design-time phases. Such isolated techniques could be insufficient in safety-critical, realistic contexts such as automotive perception systems. In this paper, we analyze design-time and run-time verification and validation techniques based on a taxonomy presented by Huang et al. We present an automotive study in which a design-time qualification workflow is proposed to complement run-time monitoring. This workflow combines a fixed safety-relevant ontology-based structured annotation system together with a synonym-based evaluation process to statistically evaluate three state-of-the-art VLMs against data from the nuScenes dataset. We observed that the proposed technique enables deterministic and repeatable quantification of the hallucinations VLMs generate in automotive perception-related tasks. The proposed workflow supports model comparison and deployment-oriented engineering decisions within the design-time verification and validation process and will contribute to a holistic verification strategy that strives towards trustworthy automotive perception systems
CommentsAccepted in ICTSS 2026 - 38th International Conference on Testing Software and Systems