基于深度学习的相机陷阱图像昆虫层级分类
Deep learning-based hierarchical insect classification using camera trap imagery
浏览论文内容
中文总结 AI 辅助
本文构建长尾昆虫数据集,提出适配可变深度层级的深度学习分类模型,结合类别平衡加权,在5级分类任务上取得80%-99%的每级别准确率,可用于昆虫生物多样性监测。
中文摘要 AI 辅助
昆虫种群衰退使可靠的生物多样性监测愈发迫切,但昆虫生物多样性监测受限于标准化数据的匮乏,以及昆虫学家手动鉴定的高成本与耗时。基于深度学习的图像分类器可处理自动化非致命相机陷阱的数据,有望变革并拓展昆虫生物多样性监测的规模,但仍存在专家标注数据集获取困难、能跨不同分类级别泛化的模型架构开发、高度不平衡数据的模型训练等挑战;层级数据还要求模型在对精细分类级别不确定时,默认输出置信度更高的粗级别预测。本文提出基于深度学习的层级分类模型应对上述挑战:首先,构建经人工整理的长尾数据集,包含约100万张昆虫图像,源自1801段相机陷阱视频,采用5级34类层级标注;其次,调整层级分类模型架构以适配5级可变深度层级,并引入类别平衡加权。该模型利用生物分类学提取粒度特定的视觉特征,生成符合层级一致性的预测,直至满足置信度阈值(T=0.6)的最深分类级别,在测试数据的5个层级上实现了80%-99%的每级别准确率。
英文摘要
Declining insect populations make reliable biodiversity monitoring increasingly urgent, yet monitoring of insect biodiversity is hampered by a lack of standardised data and by costly and time-consuming manual identification by expert entomologists. Deep learning-based image classifiers, processing data from automated non-lethal camera traps, have the potential to transform and scale insect biodiversity monitoring. However, challenges remain in acquiring expert-annotated datasets, developing model architectures that generalise well across diverse taxonomic levels and training models on highly imbalanced data. Hierarchical data also benefits from designing models that default to higher-confidence, coarser-level predictions, when uncertain about finer taxonomic levels. In this paper we address these challenges with a deep learning-based hierarchical classification model. First, we present a manually curated, long-tailed dataset of around one million images of insects, extracted from 1,801 camera-trap video recordings and annotated with a five-level, 34-class hierarchy. Further, we adapt a hierarchical classification model architecture to a five-level variable-depth hierarchy, with class-balanced weighting. Our model improves on non-hierarchical classifiers by leveraging biological taxonomy to extract granularity-specific visual features and makes hierarchy-consistent predictions to the deepest taxonomic level that meets a confidence threshold (T = 0.6). Our model achieved a per-level accuracy of 80-99% on test data, across five levels of hierarchy. Furthermore ...