arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

并非所有层都同等重要:面向可靠 CLIP 分布外检测的动态层路由

Not All Layers Are Equal: Dynamic Layer Routing for Reliable CLIP OOD Detection

Ignacio M. De la Jara, Cristian Rodriguez-Opazo, Damith Ranasinghe

arXiv 2609.20299首次发表:更新:

发表机构

University of Adelaide; Australian National University; Naval Group Pacific(阿德莱德大学; 澳大利亚国立大学; 海军集团太平洋公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出 Voyager,一种可学习的轻量级路由器,用于在 CLIP 层深度中选择稀疏的最终层锚定专家,以提升分布外检测性能,在 ImageNet-1K 上显著优于提示学习方法,且训练高效。

AI 中文摘要

跨模型层的信息聚合已被证明能改进分布外(OOD)检测。与近期工作中设计逐层信息聚合方法不同,我们探究层选择是否是一个可学习的问题。换言之,我们将问题从“如何融合层”转变为“对于给定输入,应信任哪些层”。利用一种可泛化的、弱的分布外上下文构造方法进行监督,该方法被证明比现有最先进方法的机制更有效,我们学习一个轻量级路由器,以在 CLIP 的层深度上为 OOD 检测选择一个稀疏的、以最终层为锚点的专家。在三个不同的基准上,我们证明了所提出的可学习路由方法(名为 Voyager)能改进 OOD 检测。在 ImageNet-1K 上,Voyager 实现了 18.86 的平均 FPR@95,比最强的、可比较的提示学习方法高出 8.8 个百分点。这些提升在多种监督源(包括现有最先进提示学习方法所使用的监督源)下均持续存在,这表明尽管我们的弱 OOD 监督上下文非常有效,但关键优势来自可学习的路由器组件,而非监督源。重要的是,Voyager 非常实用;路由器学习大约需要两分钟,内存占用不到 1 GB,比当前提示学习方法高效约 20 倍。匿名代码:此 https URL。

英文摘要

Information aggregation across model layers are revealed to improve OOD detection. In contrast to crafting a method for layer-wise information aggregation in recent work, we investigate if layer selection is a learnable problem. In other words, we transpose the question from how to fuse layers to one asking which layers to trust for an input. Using a generalizable, weak, out of distribution context crafting approach for supervision, shown to be more effective than state of the art methods' mechanisms, we formulate learning a lightweight router to select a sparse, final-layer-anchored expert over CLIP's layer depth for OOD detection. Across three diverse benchmarks we demonstrate our learnable routing method dubbed Voyager improves OOD detection. On ImageNet-1K, Voyager achieves an average FPR@95 of 18.86, outperforming the strongest, comparable, prompt-learning method by 8.8 points. These gains persist across multiple supervision sources, including those used by existing state-of-the-art prompt-learning methods, demonstrating that, whilst our weak OOD supervision context is highly effective, the key advantage is realized from the learnable router component rather than the supervision source. Significantly, Voyager is highly practical; router learning takes approximately two minutes using less than 1 GB of memory, making it approximately 20x more efficient than current prompt-learning approaches. Anonymized Code: https://anonymous.4open.science/r/Voyager/

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑