arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

并非所有混淆都相同:面向细粒度飞机检测的源感知不确定性诊断

Not All Confusion Is Equal: A Source-Aware Uncertainty Diagnosis for Fine-Grained Aircraft Detection

Hai Huang, Helmut Mayer

arXiv 2609.29959首次发表:更新:

发表机构

Universität der Bundeswehr München(慕尼黑联邦国防军大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出$A^2E^2$诊断工具,将细粒度飞机检测中的混淆按偶然/认知和类内/类间两轴分解为四类来源,分别测量并指导针对性干预,将混淆测量转化为可验证的、可操作的诊断。

AI 中文摘要

细粒度目标检测器通常使用混淆矩阵进行评估,混淆矩阵显示模型在何处混淆,但不显示为何混淆,也不显示混淆是否可减少。我们认为混淆可归因于不同的、可分离的来源,每种来源均可定量测量,从而将被动测量转化为可操作的指导。我们提出$A^2E^2$,一种诊断工具,将混淆来源沿两个轴分解,即$\{$偶然性,认知性$\} \ imes \{$类内,类间$\}$,形成一个$2\ imes2$的分类法,枚举来源类型。每个象限由其自身的量测量,该量在三个位置之一(输入几何、输出空间分歧和偏差参数后验)计算,因此两个认知来源通过构造而非经验相关性分离。在细粒度飞机检测上,四个象限成为四个具名来源,各自有其补救判定:亲和性(几何相似性,仅从尺寸不可减少)、异质性(几何异构子变体,指向重新标注而非更多数据)、争议性(训练不足但可学习的边界,可改进)和坍缩性(数据匮乏的类别,可减少)。在将混淆归因于特定可减少来源后,我们应用有针对性的干预,并通过实验验证其专门减少所诊断的来源,同时保持不可减少来源不变。$A^2E^2$因此将混淆测量转化为具体、可验证且可操作的“诊断”,其中相同的非对角线质量可携带相反的原因和相反的补救措施。我们还陈述了该框架的局限性,包括在该特定数据集上哪些来源仅部分可识别以及原因。

英文摘要

Fine-grained object detectors are commonly evaluated with confusion matrices, which show where the model is confused but not why, nor whether the confusion can be reduced. We argue that confusion can be attributed to distinct, separable sources, each quantitatively measurable, turning a passive measurement into actionable guidance. We present $A^2E^2$, a diagnostic tool that decomposes the sources of confusion along two axes, $\{$aleatoric, epistemic$\} \times \{$within-class, between-class$\}$, giving a $2\times2$ taxonomy that enumerates the source types. Each quadrant is measured by its own quantity, computed in one of three places (input geometry, output-space disagreement, and the bias-parameter posterior), so the two epistemic sources are separated by construction rather than by an empirical correlation. On fine-grained aircraft detection, the four quadrants become four named sources with their own remedy verdict: affinity (geometric similarity, irreducible from size alone), heterogeneity (geometrically heterogeneous sub-variants, pointing to re-labeling rather than more data), contested (an insufficiently trained but learnable boundary, improvable), and collapsed (a class starved of data, reducible). After attributing the confusion to a specific reducible source, we apply a targeted intervention and verify experimentally that it reduces the diagnosed source specifically while leaving the irreducible sources unchanged. $A^2E^2$ thus turns confusion measurement into a concrete, validatable and actionable "diagnosis" in which the same off-diagonal mass can carry opposite causes and opposite remedies. We also state this framework's limits, including which sources are only partially identifiable on this specific dataset and why.

Comments23 pages, 4 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑