发表机构
Collaborative Earth; School of Mathematics & Statistics, University of Glasgow(协作地球; 格拉斯哥大学数学与统计学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究探究地面真值不确定性对相机陷阱图像深度学习分类的影响,发现适度标签分歧训练结合ImageNet预训练可提升准确率,为整合人类与深度学习分类提供了方向。
AI 中文摘要
监督式深度学习方法可快速处理生态图像数据,但依赖成本高昂的标注流程,因此训练标签通常来自志愿者公民科学项目。然而,志愿者之间的分歧会给模型训练和验证所依赖的“地面真值”数据引入不确定性。本研究利用两个包含相机陷阱图像及对应志愿者和专家分类的数据集,探究了在更高地面真值不确定性下训练的影响。研究发现,整体测试准确率有所提升,尤其是对志愿者而言更难的图像;物种级准确率通常也有所改善,但对不同数据集的泛化能力并未提升。在ImageNet上进行预训练可增强地面真值不确定性带来的益处,还能减少所需的训练轮次;在其他相机陷阱图像上进行额外预训练可进一步降低计算成本,但无法提升准确率。在训练数据不平衡的情况下,仍观察到增加地面真值不确定性对整体准确率(尤其是困难图像)有明显益处;类别不平衡会提升常见物种的准确率,降低稀有物种的准确率,并使误分类模式更接近志愿者的错误。本研究结果对将深度学习应用于具有多个标签的各类生态图像具有重要意义,从业者可通过在训练中纳入适度的标签分歧,并使用在通用图像数据上预训练的模型,提升准确率,尤其是困难样本的准确率。除了改进公民科学衍生标签在模型训练中的应用外,本研究还为在组合工作流程中更有效地整合人类与深度学习分类提供了途径。
英文摘要
Supervised deep learning methods enable the rapid processing of ecological image data, but depend on a costly annotation process. Consequently, training labels are commonly derived from volunteer citizen science projects. However, disagreement among volunteers introduces uncertainty in the "ground truth" data that are assumed to be correct for model training and validation. Using two datasets containing camera trap images with associated volunteer and expert classifications, we investigated the effects of training under higher ground truth uncertainty. We observed improved overall test accuracy, particularly for images that were more difficult for volunteers. Species-level accuracy also generally improved, but generalisation to a different dataset did not. The benefits of ground truth uncertainty were enhanced by pre-training on ImageNet. Pre-training also reduced the number of training epochs required; further reductions in computational cost, but not gains in accuracy, resulted from additional pre-training on other camera trap images. With unbalanced training data, we still observed a clear benefit of increased ground truth uncertainty for overall accuracy, especially on difficult images. Class imbalance improved accuracy for common species, reduced rare species accuracy, and changed patterns of misclassification to more closely resemble mistakes made by volunteers. Our findings have implications for applying deep learning across ecological image types with multiple labels. Practitioners can improve accuracy, especially on difficult examples, by including moderate levels of label disagreement during training and using models pre-trained on general image data. In addition to improving the use of citizen science-derived labels in model training, our study suggests avenues for more effectively integrating human and deep learning classifications in combined workflows. (abridged)
Comments28 pages, 5 figures