生物多样性监测中的主动学习:从标签效率到可靠的生态推断
Active Learning for Biodiversity Monitoring: From Label Efficiency to Reliable Ecological Inference
浏览论文内容
中文总结 AI 辅助
针对生物多样性监测中专家标注受限的问题,综述主动学习在声学与图像模态的应用,提出兼顾训练、验证及生态推断的预算分配框架与未来路线图。
中文摘要 AI 辅助
专家标注能力的有限性是生物多样性监测中普遍存在的制约因素。被动声学记录器和相机陷阱生成数据的速度超过了专家分析它们的能力。机器学习(ML)模型可以大规模处理这些数据,但其可靠性取决于标注样本的质量、数量和覆盖范围,因此专家时间仍然是一个制约因素。主动学习(AL)通过在一个固定的标注预算下,选择那些预期最能改进模型的样本,从而缓解了这一瓶颈,已发表的证据表明,它可以减少达到目标性能所需的标签数量。然而,监测项目面临一个更广泛的问题:有限的专家预算应如何分配,才能使模型训练、验证以及基于模型输出构建的生态估计都保持可靠?由于主动学习非随机地选择样本,其标签不适用于验证、校准或阈值选择,这一矛盾很少被承认。我们综合了跨声学和图像模态的主动学习研究,并指出了差距和机遇。大多数研究在预标注的基准上使用模拟标注者评估查询策略;在真实监测工作流程中的部署很少,且集中于鸟类和鲸目动物。蝙蝠、昆虫、两栖动物和鱼类代表性不足,多模态应用在很大程度上仍未探索。评估集中于标注工作量的显著减少,通常没有随机采样基线、逐类结果或校准分析,也很少考虑验证所需的标签。我们提供了对主动学习循环的教程式处理,使这些预算决策明确化,并提出了一个支持标签高效训练、验证和可信赖的下游生态推断的主动学习方法的路线图。
英文摘要
Limited expert annotation capacity is a pervasive constraint in biodiversity monitoring. Passive acoustic recorders and camera traps generate data faster than experts can analyse them. Machine learning (ML) models can process these data at scale, but their reliability depends on the quality, quantity, and coverage of labelled samples, so expert time remains a constraint. Active learning (AL) eases this bottleneck by selecting, under a fixed annotation budget, the samples expected to improve a model most, and published evidence shows it can reduce the labels needed to reach a target performance. Monitoring programmes, however, face a broader question: how should a limited expert budget be divided so that model training, validation, and the ecological estimates built on model outputs all remain reliable? Because AL selects samples non-randomly, its labels are unsuitable for validation, calibration, or threshold selection, a tension rarely acknowledged. We synthesise AL research across acoustic and image modalities and identify gaps and opportunities. Most studies evaluate query strategies on pre-labelled benchmarks with simulated annotators; deployments in real monitoring workflows are rare and concentrate on birds and cetaceans. Bats, insects, amphibians, and fish are underrepresented, and multimodal applications remain largely unexplored. Evaluation centres on headline reductions in annotation effort, often without random-sampling baselines, per-class results, or calibration analysis, and rarely accounts for the labels required for validation. We provide a tutorial treatment of the AL loop that makes these budget decisions explicit, and a roadmap towards AL methods that support label-efficient training, validation, and trustworthy downstream ecological inference.
发表机构
- University of Amsterdam(阿姆斯特丹大学)
- Tampere University(坦佩雷大学)
- Leiden University(莱顿大学)
- Naturalis Biodiversity Centre(自然生物多样性中心)
机构由 AI 辅助整理,请以论文原文为准。