VD-DeepStack:弥合视觉比较与语言推理的少样本异常检测
VD-DeepStack: Bridging Visual Comparison and Language Reasoning for Few-Shot Anomaly Detection
另 1 家 · 查看机构详情
- Fudan University(复旦大学)
- Tongji University(同济大学)
- University of Southern California(南加州大学)
- Fuzhou University(福州大学)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
提出VD-DeepStack,通过融合DINO特征并构建密集差异证据注入解码器,弥合视觉比较与语言推理的差距,在工业和医学基准上显著提升少样本异常检测性能。
中文摘要 AI 辅助
少样本视觉异常检测本质上是一项视觉比较任务,需要针对正常参考对查询进行细粒度检查。许多基于大型视觉语言模型(LVLMs)的最新方法强调通过语言思维链进行比较推理。然而,离散的抽象描述可能无法充分代表密集的细粒度视觉差异,在视觉比较及其语言表达之间留下差距。为解决这一差距,我们提出了视觉差异深度堆叠(VD-DeepStack),该方法明确地将语言推理条件建立在查询-参考视觉差异之上。具体而言,我们将DINO特征与LVLM视觉层次结构融合,以增强细粒度表示,然后从查询特征与软匹配参考特征之间的残差构建密集差异证据。差异证据路径在多个解码器深度将空间加权差异向量注入查询图像状态,而辅助视觉上下文路径提供细粒度外观信息以支持其解释。在4个工业和2个医学异常基准上的实验表明,与依赖文本比较推理的基线相比,少样本异常检测有显著改进。这些结果支持通过比较表示的联合设计及其在解码器中的集成来弥合视觉比较-推理差距。代码将在录用后发布。
英文摘要
Few-shot visual anomaly detection is fundamentally a visual comparison task, requiring fine-grained inspection of a query against normal references. Many recent methods based on large vision-language models (LVLMs) emphasize comparative reasoning through language chain-of-thought. Yet discrete, abstract descriptions may underrepresent dense, fine-grained visual differences, leaving a gap between visual comparison and its expression in language. To address this gap, we propose Visual Difference DeepStack (VD-DeepStack), which explicitly conditions language reasoning on query-reference visual differences. Specifically, we fuse DINO features with the LVLM visual hierarchy to strengthen fine-grained representations, then construct dense difference evidence from residuals between query features and softly matched reference features. The difference-evidence path injects spatially weighted difference vectors into query-image states at multiple decoder depths, while an auxiliary visual-context path provides fine-grained appearance information to support their interpretation. Experiments on 4 industrial and 2 medical anomaly benchmarks demonstrate substantial improvements in few-shot anomaly detection over baselines relying on textual comparative reasoning. These results support mitigating the visual comparison-reasoning gap through the joint design of comparison representations and their integration into the decoder. Code will be released upon acceptance.