视觉-语言表示学习中的动态分布感知不确定性跟踪
Dynamic Distribution-Aware Uncertainty Tracking in Vision-Language Representation Learning
浏览论文内容
中文总结 AI 辅助
针对视觉-语言模型事后不确定性量化方法忽略测试分布动态性的问题,提出DDA-UQ框架,通过高斯混合模型建模嵌入空间,实现动态不确定性估计,性能优于现有最优方法。
中文摘要 AI 辅助
不确定性量化(UQ)旨在衡量模型预测的可靠性,是将视觉-语言模型(VLMs)部署到安全关键场景的关键保障。事后方法因轻量特性被广泛采用,通过可学习模块或归纳式总结将VLMs的输出映射为不确定性度量,但这类方法本质上仍局限于拟合源域的失败模式,忽略了测试分布的动态特性。为应对该挑战,我们提出动态分布感知不确定性量化框架(DDA-UQ),将范式从静态映射转向动态分布感知过程。训练期间,我们利用高斯混合模型对VLMs的嵌入空间建模并提取分布证据,从而动态推导不确定性估计;推理阶段,该设计可对数据分布的变化做出动态响应。大量实验表明,我们的方法显著优于现有最优方法。
英文摘要
Uncertainty Quantification (UQ) aims to measure the reliability of model predictions, serving as a critical safeguard for deploying Vision-Language Models (VLMs) in safety-critical scenarios. Post-hoc approaches are widely adopted due to their lightweight nature, mapping the outputs of VLMs to uncertainty measures through learnable modules or inductive summarization. However, Post-hoc approaches remain inherently confined to fitting the failure patterns of the source domain, ignoring the dynamic nature of test distributions. To address this challenge, we propose a Dynamic Distribution-Aware Uncertainty Quantification framework (DDA-UQ) that shifts the paradigm from static mapping to a dynamic distribution-aware process. During training, we leverage a Gaussian Mixture Model to model the VVLMs'embedding space and extract distributional evidence, thereby dynamically deriving uncertainty estimates. During inference, the design dynamically responds to changes in the data distribution. Extensive experiments demonstrate that our approach significantly outperforms state-of-the-art methods.
发表机构
- State Key Laboratory of Novel Software Technology, Nanjing University(南京大学计算机软件新技术国家重点实验室)
- ByteDance Inc(字节跳动公司)
机构由 AI 辅助整理,请以论文原文为准。