arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

公共葡萄病害数据集中的捷径学习:标注粒度作为调节因素而非原因

Shortcut Learning in a Public Grape Disease Dataset: Annotation Granularity as a Modulator, Not a Cause

Pushuo Wang

arXiv 2608.20663首次发表:更新:

AI 中文总结

研究发现公共葡萄病害数据集的标注粒度是捷径学习的调节因素而非原因,提出无需图像或训练的粒度筛选统计量,且空中病斑级检测光学上不可行。

AI 中文摘要

农业病害检测的公共数据集通常通过报告的指标判断是否适用,但这些指标无法说明标注方案是否内部一致。在一个包含3288张图像、11995个边界框、6个类别的公共葡萄病害数据集上,改变模型容量、输入分辨率和检测范式得到的测试集mAP50范围与种子间噪声相当,瓶颈在于所有5种架构中的小目标。该发现源于数据层面:其中1个类别以整叶级别标注(边界框中位数面积为图像的43.16%),其余5个类别以病斑级别标注。在5156张不含葡萄的跨物种图像中,65.7%的假阳性边界框属于该类别,其假阳性数量相对于训练标注中的占比高出13.41倍。反事实重训练证实了粒度对捷径幅度的因果效应:仅缩小该类别边界框可将其跨物种假阳性降低66%,安慰剂对照确认该效应仅针对被操纵的类别。采用预先注册的标准进行反向操纵则得到阴性结果:将最精细的类别粗化为整叶级别(边界框占比从0.57%变为40.37%),其边界框数量和标注占比匹配,且分布内AP更高,但其跨物种假阳性仍为0个,而未被操纵的原始类别占假阳性的50.0%。因此,标注粒度是该捷径的调节因素而非原因:它可放大或衰减已存在的陷阱,但无法创造陷阱,且修复该陷阱的方法仍未知。我们还提出了一种无需图像或训练的粒度筛选统计量,并表明空中病斑级检测在光学上不可行。这种失败模式对分布内评估不可见。

英文摘要

Public datasets for agricultural disease detection are usually judged fit for use from reported metrics, which say nothing about whether the annotation scheme is internally consistent. On one public grape disease dataset (3288 images, 11995 boxes, 6 classes), varying model capacity, input resolution and detection paradigm yields a test-set mAP50 range comparable to seed-to-seed noise, with the bottleneck at small objects across all five architectures. The finding lies on the data side: one class is annotated at whole-leaf level (median box area 43.16% of the image) while the other five are annotated at lesion level. On 5156 cross-species images containing no grape, 65.7% of the false-positive boxes fall into that one class, an over-representation of 13.41x relative to its share of the training annotations. Counterfactual retraining establishes a causal effect of granularity on the magnitude of the shortcut: shrinking only that class's boxes cuts its cross-species false positives by 66%, and a placebo control confirms the effect is specific to the manipulated class. A manipulation in the opposite direction, with criteria registered in advance, returns a negative result: coarsening the finest class to whole-leaf level (0.57% to 40.37%), matched in box count and share of annotations and with higher in-distribution AP, still leaves its cross-species false positives at zero boxes, while the unmanipulated original class holds 50.0% of them. Annotation granularity is therefore a modulator of this shortcut, not its cause: it can amplify or attenuate a sink that already exists, but cannot create one, and what fixes the destination remains open. We also give a granularity screening statistic requiring neither images nor training, and show airborne lesion-level detection to be optically out of reach. The failure mode is invisible to in-distribution evaluation.

Comments30 pages, 3 figures, 17 tables. Code and evaluation artifacts: https://github.com/nck9343-a11y/crop-detect

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑