arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

在分布偏移下,边际覆盖率能否保证零样本视觉语言模型(VLM)的类条件安全性?

Does Marginal Coverage Guarantee Class-Conditional Safety for Zero-Shot VLMs Under Shift?

Jai Kumar Sharma, Amartya Dutta

arXiv 2608.19376首次发表:更新:

发表机构

Virginia Tech(弗吉尼亚理工大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对CLIP等零样本VLM,发现边际共形覆盖率无法保证类条件安全性,源域校准无法迁移,目标域类别校准可改善尾部但需类别标签,指出其仅为平均可靠性统计量。

AI 中文摘要

分裂共形预测在可交换性下提供边际覆盖率,正被越来越多地用作零样本视觉语言模型(VLM)的弃权(不执行)层。我们针对CLIP、OpenCLIP和SigLIP在ImageNet及非ImageNet设置下的部署偏移情况,对这一做法进行审计。边际覆盖率可保持相对较高,而类条件尾部覆盖率却会崩溃:在ImageNet-Sketch上,尽管边际覆盖率约为0.86,但最差类覆盖率降至≈0,且有10-12%的类别低于有限样本零基线。该失败情况与目标域类别准确率一致,但无法通过我们测试的源域诊断指标预测。源域Mondrian校准可改善分布内尾部,但无法迁移;聚类共形和Conf-OT可改善边际或平均指标,但无法恢复最差类尾部。目标域类别校准可大幅提升尾部性能,但需要每个类别的标签,且设置规模密集。我们还发现2-3倍的跨系列效率差距,并表明原生SigLIP的sigmoid分数消除了APS的概率质量解释。这些发现在测试的模型规模、预训练语料库、提示、未覆盖率α及偏移的非ImageNet设置中均成立。因此,边际共形覆盖率应被视为平均可靠性统计量,而非类尾部的安全保证。

英文摘要

Split-conformal prediction provides marginal coverage under exchangeability and is increasingly used as an abstention layer for zero-shot vision-language models (VLMs). We audit this practice under deployment shift for CLIP, OpenCLIP, and SigLIP across ImageNet and non-ImageNet settings. Marginal coverage can remain relatively high while class-conditional tail coverage collapses: on ImageNet-Sketch, worst-class coverage falls to $\approx 0$ and 10-12% of classes lie below a finite-sample null floor, despite marginal coverage of about 0.86. The failure is aligned with target-domain class accuracy but is not predicted by the source-domain diagnostics we test. Source-side Mondrian calibration improves the in-distribution tail but does not transfer, while clustered conformal and Conf-OT improve marginal or average metrics without recovering the worst-class tail. Target-side class calibration substantially lifts the tail, but requires labels for every class and remains set-size-intensive. We further identify a 2-3$\times$ cross-family efficiency gap and show that native SigLIP sigmoid scores remove APS's probability-mass interpretation. The findings persist across the tested model scale, pretraining corpus, prompt, miscoverage level $α$, and shifted non-ImageNet settings. Marginal conformal coverage should therefore be treated as an average reliability statistic, not as a safety guarantee for the class tail.

CommentsAccepted at the ECCV 2026 Workshop on Uncertainty Quantification for Computer Vision (UNCV). 34 pages (16 main + 18 supplementary), 10 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑