arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.35420cs.CV

用于野生动物保护的相机陷阱图像自动物种识别

Automated Species Identification in Camera Trap Images for Wildlife Conservation

  • Brac University(布拉克大学)

机构由 AI 辅助整理,请以论文原文为准。

Nowshin Amin, Nafisa Tabassum Oyshi, Tahmid Abrar Zidan, Miftaun Noor, Md. Abrar Rahman Shafin

AI总结:

本研究提出一种集成自注意力机制和Swin-BiFPN骨干的Faster RCNN框架,结合LLaVA多模态模型,提升低对比度陷阱图像中小型动物的检测精度,并实现零样本检测。

AI中文摘要:

野生动物保护涉及保护、保存和管理野生动物物种及其栖息地。随着当今人类发展、气候变化和其他不可持续做法的快速推进,野生动物保护的需求日益增强。尽管使用深度学习模型在物种识别方面取得了显著进展,但由于特征提取能力有限,在低对比度陷阱图像中有效检测小型动物仍然面临重大挑战。本论文提出了一种集成自注意力机制的新型端到端框架,以解决这些局限性。所提出的架构包括一个集成在Faster RCNN检测网络中的Swin-BiFPN骨干网络,并配有一个由LLaVA v1.5(13B)多模态大语言模型驱动的视觉语义提取模块。该检测框架能够在具有挑战性的陷阱图像中提取关键特征,展现出持续的高性能和强大的泛化能力。此外,视觉语义提取模块提供了零样本检测能力,并提供了关于动物行为的有价值见解和新出现的线索,进一步支持保护工作。多模态大语言模型评估使用了传统NLP指标(精确率、召回率、F1和SBERT相似度)以及基于LLM的评判者(GPT-4.1和GROK 3.0)的主观评分,对五个多模态大语言模型进行了评估,证明了该模型在视觉描述生成方面的强大性能。所提出的框架提高了低对比度陷阱图像和小型动物的检测准确性,同时展示了利用多模态大语言模型的零样本检测能力。

英文摘要:

Wildlife conservation involves protecting, preserving, and managing wildlife species and their habitats. With today's rapid pace of human development, climate change, and other unsustainable practices, the need for wildlife conservation has heightened. Despite significant progress in species identification using deep-learning models, significant challenges still remain in effectively detecting small animals in low-contrast trap images due to limited feature extraction capabilities. This thesis presents a novel end-to-end framework integrating a self-attention mechanism to address these limitations. The proposed architecture involves a Swin-BiFPN backbone integrated in a Faster RCNN detection network, coupled with a visual semantic extraction module driven by the LLaVA v1.5 (13B) multimodal large language model. The detection framework, capable of extracting crucial features in challenging trap images, demonstrates consistently high results and robust generalization capabilities. Furthermore, the visual semantic extraction module provides zero-shot detection capability, as well as providing valuable insights and emergent cues of the animal's behavior, further supporting the conservation effort. The MLLM evaluation was conducted using both traditional NLP metrics (precision, recall, F1, and SBERT similarity) and subjective scoring by LLM-based judges (GPT-4.1 and GROK 3.0), across five MLLMs, demonstrating the model's strong performance in visual description generation. The proposed framework improves detection accuracy across low-contrast trap images and small animals while also demonstrating zero-shot detection capability leveraging the MLLM.

补充信息

↑