视觉Transformer与卷积神经网络在新几内亚兰科植物细粒度属级识别中的对比:物种丰富但数据匮乏的植物区系上的受控基准研究
Vision Transformers versus convolutional neural networks for fine-grained orchid genus identification in a species-rich, data-poor flora: a controlled benchmark on the Orchidaceae of New Guinea
浏览论文内容
中文总结 AI 辅助
针对新几内亚兰花物种丰富但数据匮乏的现状,本研究构建两阶段细粒度识别系统,比较四种预训练骨干,发现自监督ViT(DINOv2)结合嵌入检索表现最优,并发布开放网络应用作为实用模板。
中文摘要 AI 辅助
新几内亚拥有世界上最丰富的岛屿植物区系(约2,856种兰花),但大多数物种仅有少量照片可供使用,远少于直接进行物种级分类所需的数量。在这样物种丰富但数据匮乏的植物区系中,需要细粒度识别方法,但目前尚不清楚哪种骨干架构和预训练策略最能支持这些方法。我们构建了一个两阶段系统,首先预测查询照片的属,然后使用FAISS检索候选物种的视觉相似参考图像。我们比较了四种预训练骨干——两种视觉Transformer(ViT;DINOv2、BioCLIP 2)和两种卷积神经网络(ConvNeXt V2-L、EfficientNetV2-L)——在固定的、按物种分层的16,701张照片分区上,以相同协议进行微调,涵盖120个属和1,350个物种,评估了准确性、校准、错误结构、物种检索以及新属的开放集检测。DINOv2取得了最佳的属级性能(宏平均top-1为66.9%,95%置信区间63.7-70.6;全局top-1为88.9%);两种ViT均优于两种CNN,且通用自监督预训练(DINOv2)比领域匹配的生物预训练(BioCLIP 2)在宏平均top-1上高出7.1个百分点。错误集中在两个作为错误吸引子的丰富属上。DINOv2嵌入实现了物种Recall@5为86.6%和属级Recall@5为98.7%;温度缩放将每个骨干的期望校准误差降低至约0.03;基于距离的开放集门控标记了未见过的属(平均AUROC为0.958)。自监督视觉Transformer骨干结合嵌入检索是物种丰富但数据匮乏植物区系中细粒度识别的有效且可部署的策略。该系统已作为开放网络应用程序(新几内亚兰花识别器)发布,为其他高度多样化、记录不足的分类群提供了实用模板。
英文摘要
New Guinea is the world's richest island flora (~2,856 orchid species), yet most species are represented by only a handful of photographs, far fewer than direct species-level classification requires. Methods for fine-grained identification in such species-rich, data-poor floras are needed, and it remains unclear which backbone architecture and pretraining strategy best support them. We built a two-stage system that first predicts the genus of a query photograph, then retrieves visually similar reference images of candidate species using FAISS. We compared four pretrained backbones -- two Vision Transformers (ViTs; DINOv2, BioCLIP 2) and two CNNs (ConvNeXt V2-L, EfficientNetV2-L) -- fine-tuned under an identical protocol on a fixed, species-stratified partition of 16,701 photographs spanning 120 genera and 1,350 species, assessing accuracy, calibration, error structure, species retrieval, and open-set detection of novel genera. DINOv2 attained the best genus performance (macro top-1 66.9%, 95% CI 63.7-70.6; global top-1 88.9%); both ViTs outranked both CNNs, and general-purpose self-supervised pretraining (DINOv2) outperformed domain-matched biological pretraining (BioCLIP 2) by 7.1 points of macro top-1. Errors concentrated on two abundant genera acting as error attractors. DINOv2 embeddings achieved species Recall@5 of 86.6% and genus Recall@5 of 98.7%; temperature scaling reduced every backbone's Expected Calibration Error to about 0.03; and a distance-based open-set gate flagged unseen genera (mean AUROC 0.958). A self-supervised Vision-Transformer backbone combined with embedding retrieval is an effective, deployable strategy for fine-grained identification in species-rich, data-poor floras. The system is released as an open web application (the New Guinea Orchid Identifier), offering a practical template for other hyperdiverse, under-documented taxa.
发表机构
- James Cook University(詹姆斯库克大学)
- Southwest Papua Natural Resources Conservation Agency, Ministry of Forestry, Indonesia(印度尼西亚林业部西南巴布亚自然资源保护局)
- National Research and Innovation Agency (BRIN), Republic of Indonesia(印度尼西亚共和国国家研究与创新局(BRIN))
- Royal Botanic Gardens, Kew(英国皇家植物园邱园)
- Australian Tropical Herbarium(澳大利亚热带植物标本馆)
- Commonwealth Industrial and Scientific Research Organisation (CSIRO)(澳大利亚联邦科学与工业研究组织(CSIRO))
- Australian National Herbarium(澳大利亚国家植物标本馆)
机构由 AI 辅助整理,请以论文原文为准。