arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.32567cs.CVcs.CL

零训练LLM+OVOD流水线中的归因差距:CAAP--SNAP差异的细粒度分析

Attribution Gaps in Zero-Training LLM+OVOD Pipelines: A Fine-Grained Analysis of the CAAP--SNAP Discrepancy

Yu-Feng Yen

AI总结:

本文在完整COCO-Val上分析零训练LLM+OVOD流水线中CAAP与SNAP的差异,发现差距源于封闭类别标注限制而非真实定位失败,词汇新颖性导致精度显著下降,且该现象具有跨模型通用性。

AI中文摘要:

LAOD及类似的零训练LLM+开放词汇检测器(OVOD)流水线分别对两项指标进行评分:类别无关的定位精度(CAAP)和语义命名精度(SNAP)。这两项指标持续存在分歧,但此前无人探究其原因。本文在完整的5,000张图像的COCO-Val分割(27,273个检测结果)上提出并回答了这一问题,而非原始工作所评估的小型子集。物体视觉复杂度并非驱动因素——小型和遮挡物体在定位上甚至优于大型物体。词汇新颖性才是关键:一旦LLM的措辞超出检测器的原生类别集合,定位精度从80.9%下降至31.6%。然而,这种下降并非均匀分布于不熟悉的措辞中。几乎所有下降都源于新颖措辞实际命名了与COCO标注不同的物体(真正的同义词仍达到89.3%的精度;语义无关的“噪声”标签则仅为12.0%)。对进一步失败子集的深入分析揭示了类似的故事:78-88%看似完全定位失败的情况,实际上是模型正确找到了COCO非穷尽式80类方案从未标注的真实物体,而非幻觉。更换检测器骨干网络(将YOLO-World替换为Grounding DINO)或LLM(将Gemma-3替换为Qwen2.5-VL),该效应及其大致幅度均保持稳定,这表明这似乎是流水线家族的普遍属性,而非某一模型配对的特性。最终结论是,大部分表面上的CAAP--SNAP差距可追溯至封闭类别标注的限制,而非真实的接地失败,这对于我们如何检测幻觉、分析失败模式以及设计面向开放世界工作的接地多模态系统的评估具有重要意义。

英文摘要:

LAOD and similar zero-training LLM+open-vocabulary-detector (OVOD) pipelines score two things separately: class-agnostic localization accuracy (CAAP) and semantic naming accuracy (SNAP). The two consistently diverge, and nobody has asked why. This paper asks why, on the full 5,000-image COCO-Val split (27,273 detections) rather than the small subset the original work evaluated on. Object visual complexity turns out not to be the driver -- small and occluded objects are, if anything, localized better than large ones. Vocabulary novelty is: once the LLM's wording falls outside the detector's native category set, localization accuracy falls from 80.9% to 31.6%. That drop is not spread evenly across unfamiliar phrasing, though. Almost all of it comes from cases where the novel wording actually names a different object than the one COCO annotated (true synonyms still score 89.3%; semantically unrelated "noise" labels score 12.0%). A closer look at a further failure subset tells a similar story: 78-88% of what looks like complete localization failure is really the model correctly finding a real object that COCO's non-exhaustive 80-category scheme simply never labeled, not hallucination. Swap the detector backbone (YOLO-World for Grounding DINO) or the LLM (Gemma-3 for Qwen2.5-VL) and both the effect and its rough size hold up, so this looks like a general property of the pipeline family rather than a quirk of one model pairing. The upshot is that a large share of the apparent CAAP--SNAP gap traces back to closed-category annotation limits rather than a real grounding failure, which matters for how we detect hallucination, analyze failure modes, and design evaluation for grounded multimodal systems meant to work in the open world.

补充信息

↑