AI 中文总结
本文通过审计证明零样本OOD检测器排名无法跨域迁移,提出CEG封装器,可降低检测器敏感性并提升GL-MCM与MCM的族平衡FPR95。
AI 中文摘要
为新部署场景选择零样本分布外(OOD)检测器时,通常以基准排名为依据,隐含假设排名最高的检测器可跨域迁移。本文证明该假设不成立:通过在17个分布内数据集、3个视觉语言模型(VLM)和7种代表性零样本OOD检测器间开展可控可移植性审计,发现检测器排名会随部署场景反转,且每个检测器在至少一个域上的FPR95(95%真阳性率下的假阳性率)均超过80%,最优检测器取决于分布内数据及底层VLM。研究将排名反转归因于视觉语言logits中的互补证据通道:无语料检测器依赖绝对匹配水平与相对或空间锐度的不同组合,而基于WordNet的方法还依赖外部语义覆盖。一简单命题表明,水平与锐度无法从彼此中通用恢复,解释了为何无单一检测器可跨部署场景可靠迁移。受此诊断启发,本文提出互补证据防护器(CEG),一种检测器无关的封装器,仅利用经验分布内百分位数,通过基础检测器、水平与锐度的非补偿性融合保留互补证据;用熵、logit方差或随机噪声替代这些通道的对照实验未复现增益。在无OOD样本、辅助语料或学习型融合的情况下,CEG可降低检测器敏感性,将GL-MCM的族平衡FPR95从38.1降至28.8,MCM的族平衡FPR95从42.6降至30.5。
英文摘要
Selecting a zero-shot out-of-distribution (OOD) detector for a new deployment is typically based on benchmark rankings, implicitly assuming that the highest-ranked detector will transfer across domains. We show that this assumption does not hold. Through a controlled portability audit across seventeen in-distribution datasets, three vision-language models, and seven representative zero-shot OOD detectors, we find that detector rankings reverse across deployments, every detector exceeds $80\%$ FPR95 on at least one domain, and the preferred detector depends on both the in-distribution data and the underlying VLM. We trace these reversals to complementary evidence channels in vision-language logits. Corpus-free detectors rely on different combinations of absolute match level and relative or spatial sharpness, while WordNet-based methods additionally depend on external semantic coverage. A simple proposition shows that level and sharpness cannot generally be recovered from one another, explaining why no single detector transfers reliably across deployments. Motivated by this diagnosis, we introduce the Complementary Evidence Guard (CEG), a detector-agnostic wrapper that preserves complementary evidence through a non-compensatory fusion of the base detector, level, and sharpness using only empirical in-distribution percentiles. Controls replacing these channels with entropy, logit variance, or random noise do not reproduce the gains. Without OOD samples, auxiliary corpora, or learned fusion, CEG reduces detector sensitivity and improves GL-MCM from $38.1$ to $28.8$ and MCM from $42.6$ to $30.5$ family-balanced FPR95.