arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

多少个类别才够?少样本异常阈值的无分布认证界限

How Many Categories Are Enough? Distribution-Free Certification Limits for Few-Shot Anomaly Thresholds

Gia Huy Thai, Nguyen Thai Anh

arXiv 2610.00236首次发表:更新:

发表机构

University of Science, VNU-HCM; Van Lang University(越南国立大学胡志明市理科大学; 万朗大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究推导少样本异常检测阈值认证所需类别数量的理论下界,提出CRESS协议,证明仅靠少量源类别无法支持可迁移的可靠性声明。

AI 中文摘要

少样本异常检测器通过排序指标进行评估,然而部署时需要一个具有受控虚警率(FAR)的报警阈值。我们探究需要多少正常证据(以图像或类别为单位)才能为未见类别认证这样的阈值。使用冻结的DINOv2主成分分析(PCA)残差排序器,在15个MVTec和12个VisA类别上,涵盖四种损坏类型,我们表明仅使用目标类的留一图像(LOIO)校准受分辨率限制且对偏移敏感:秩值不能低于$1/(k+1)$,在可达到的水平$\alpha=0.20$下,当$k=4$时,高斯损坏的MVTec上经验FAR达到0.341,是名义水平的1.7倍。随后推导了类别数量可行性演算:即使类别损失全为零且不计算多重性,任何确定性的、普遍有效的、无分布的95%置信上界(UCB)在$\alpha=0.20$、$0.10$和$0.05$下分别至少需要14、29和59个独立同分布(iid)类别抽样;这些数量是必要条件但非充分条件。跨类别可靠性估计与源支持(CRESS)协议将源类别划分为不相交的参考、提议和认证角色。仅使用三个或四个认证类别,所有960种冻结配置都返回故障关闭阈值$\tau^\star=0$,最小的类别级UCB为0.950。对同一档案的图像单元分析在36.7%至60.3%的目标单元中选择正阈值;这些界限适用于选定的源混合,而非新类别抽样的边际风险。我们的贡献是定量可行性边界和估计目标感知协议,明确说明源证据何时能(以及何时不能)支持可迁移的可靠性声明。

英文摘要

Few-shot anomaly detectors are judged by ranking metrics, yet deployment requires an alarm threshold with a controlled false-alarm rate (FAR). We ask how much normal evidence, in images or category units, is needed to certify such a threshold for an unseen category. Using a frozen DINOv2 principal component analysis (PCA) residual ranker on 15 MVTec and 12 VisA categories under four corruption types, we show that target-only leave-one-image-out (LOIO) calibration is resolution-limited and shift-fragile: rank values cannot fall below $1/(k+1)$, and at the attainable level $α=0.20$, empirical FAR reaches 0.341 on Gaussian-corrupted MVTec at $k=4$, 1.7 times the nominal level. A category-count feasibility calculus is then derived: even with all-zero category losses and no multiplicity charged, any deterministic, uniformly valid, distribution-free 95% upper confidence bound (UCB) requires at least 14, 29, and 59 independent and identically distributed (iid) category draws at $α=0.20$, $0.10$, and $0.05$; these counts are necessary but not sufficient. The Cross-category Reliability Estimation with Source Support (CRESS) protocol splits source categories into disjoint reference, proposal, and certification roles. With only three or four certification categories, all 960 frozen configurations return the fail-closed threshold $τ^\star=0$, and the smallest category-level UCB is 0.950. Image-unit analyses of the same archive select positive thresholds in 36.7% to 60.3% of target cells; these bounds hold for the selected source mixture, not for the marginal risk of a new-category draw. The contribution is a quantitative feasibility boundary and an estimand-aware protocol specifying when source evidence can, and cannot, support a transferable reliability claim.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑