arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

NEEDL-Bench:用于显微镜图像中瑞士针叶枯病和气孔检测的数据集

NEEDL-Bench: Dataset for Swiss Needle Cast and Stomata Detection in Microscopy Images

Benjamin Blake, Declan McIntosh, Jürgen Ehlting, Nicolas Feau, Joey B. Tanney, Alexandra Branzan Albu

arXiv 2607.12076首次发表:更新:

AI 中文总结

针对花旗松瑞士针叶枯病检测缺乏数据集的问题,提出NEEDL-Bench数据集,含3250张带注释图像,展示了模糊等多种挑战性特征。给出两种评估划分,评估多种检测方法,最大F1分数0.8479,发现改进非源于规模法则,而是特定领域归纳偏差。

AI 中文摘要

我们提出了NEEDL-Bench,这是一个针对花旗松的瑞士针叶枯病(SNC)的显微镜检测基准。花旗松作为软木木材资源,是具有重要生态和经济意义的关键物种,SNC通过在针叶的气体交换孔(气孔)中形成有性生殖结构(假囊壳)来影响生产力,从而阻碍气体交换并损害针叶功能。尽管计算机视觉有望标准化并切实扩大严重程度测量,但目前尚无用于自动检测这些结构的数据集。为此,我们展示了NEEDL-Bench,这是一个包含来自1082枚花旗松针叶的3250张带注释图像的数据集,同时为关键点和边界框检测器提供注释。该数据集具有挑战性,包括模糊、物体对比度差、感兴趣的小物体和遮挡等特征。为了更好地捕捉数据的标称分布和罕见结构的全范围,我们提出了两种不同的评估划分:从收集的图像中随机采样或顺序采样以最大化结构多样性。我们评估了多种流行的关键点和边界框检测方法作为基线,观察到最大F1分数为0.8479,这表明未来在这个问题上的发展有显著潜力。此外,我们发现较大的模型在该数据集上通常不会表现出相应的性能提升,这表明该问题的改进并非来自规模法则,而是来自特定领域的归纳偏差。

英文摘要

We present NEEDL-Bench, a microscopy detection benchmark for Swiss Needle Cast (SNC), a fungal disease of Douglas-fir trees. Douglas-fir is a keystone species of major ecological and economic importance as a softwood timber resource, and SNC affects productivity by forming sexual reproductive structures (pseudothecia) that emerge through the gas exchange pores (stomata) of the needles, thereby blocking gas exchange and compromising needle function. To date, there is no dataset for automatic computer vision detection of these structures, despite computer vision being well poised to standardize and viably scale severity measurements. To address this, we present NEEDL-Bench, a dataset of 3250 annotated images from 1082 Douglas-fir needles, annotated for both keypoints and bounding-box detectors. This dataset exhibits a challenging collection of features, including blur, poor object contrast, small objects of interest, and occlusions. To better capture both the nominal distribution of the data and the full breadth of rare structures, we present two distinct evaluation splits: either random sampling from the collected images or sequential sampling to maximize structural diversity. We evaluate multiple popular keypoint and bounding box methods for detection on this dataset as a baseline and observe a maximum F1 score of 0.8479, suggesting significant potential for gains from future development on this problem. Further, we find that larger models generally do not show commensurate gains in performance on this dataset, indicating that improvements on this problem will not come from scaling laws but rather from domain-specific inductive biases.

Comments8 Pages, Published at CRV2026

DOI:10.21428/d82e957c.c5e5ca8e

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑