面向边缘设备的猛禽物种轻量级图像分类:通过视频帧提取、知识蒸馏和TensorRT部署实现稀有物种数据集扩展
Lightweight Image Classification of Raptor Species for Edge Devices: Rare-Species Dataset Expansion via Video Frame Extraction, Knowledge Distillation, and TensorRT Deployment
浏览论文内容
中文总结 AI 辅助
该研究通过视频帧提取扩展稀有猛禽数据集,经知识蒸馏得到轻量级模型,部署后实现边缘设备上的高效分类,缓解近缘物种混淆,为风力涡轮机碰撞 mitigation 提供支持。
中文摘要 AI 辅助
本研究针对风力涡轮机碰撞 mitigation 需求,探索面向实时边缘部署的猛禽物种轻量级分类方案。以参数量达3.04亿的DINOv2-L作为教师模型,对MobileNetV4、ViT-Small、EfficientNet-B0三款轻量级学生模型进行知识蒸馏。为降低近缘物种间的分类混淆,研究通过视频帧提取将数据集规模扩展至12519张图像,其中虎头海雕(Steller's Sea Eagle)图像从463张增至2050张。采用按视频及源图像级别划分样本的分组拆分策略以缓解该粒度下的源数据泄露问题,在5组蒸馏随机种子下,三款学生模型集成的宏召回率达0.935±0.004;在传统图像级别拆分下,宏召回率达0.955,且参数量仅为教师模型的约八分之一,保留了教师模型97.5%的宏召回率。在与训练图像不重叠的1258张图像子集上,白尾海雕(White-tailed Eagle)的召回率提升最高达38.6个百分点,其被误分类为虎头海雕的错误率从61%降至15%。在NVIDIA Jetson Orin Nano上,将EfficientNet-B0以TensorRT FP16模式部署后,含主机-设备数据传输在内的单张图像推理耗时3.19毫秒,推理速度达313张/秒,与FP32模式的argmax一致性达99.95%。在5组受控随机种子对比中,知识蒸馏(对比仅使用交叉熵损失)或教师模型从DINOv2-L更换为DINOv3-L均未在集成级别产生明显性能提升,主要增益来自数据集扩展与教师模型微调。
英文摘要
We investigate lightweight raptor-species classification for real-time edge deployment in wind-turbine collision mitigation. Using DINOv2-L (304M parameters) as a teacher, we distilled three lightweight students (MobileNetV4, ViT-Small, and EfficientNet-B0). To reduce confusion between closely related species, we expanded the dataset to 12,519 images, including an increase in Steller's Sea Eagle images from 463 to 2,050 via video-frame extraction. Under a group split that separates samples at the video- and source-image level to mitigate source leakage at that granularity, the three-student ensemble achieved a macro recall of 0.935 +/- 0.004 over five distillation seeds (0.955 on a conventional image-level split, retaining 97.5% of the teacher's macro recall) with roughly one-eighth as many parameters. On a subset of 1,258 images disjoint from the former training images, White-tailed Eagle recall improved by up to 38.6 percentage points, while the rate at which it was misclassified as the Steller's Sea Eagle decreased from 61% to 15% of errors. TensorRT FP16 deployment of EfficientNet-B0 on an NVIDIA Jetson Orin Nano achieved 3.19 ms/image including host-device transfer (313 images/s), with 99.95% argmax agreement with FP32. In five-seed controlled comparisons, neither distillation (versus CE-only) nor the change from a DINOv2-L to a DINOv3-L teacher yielded a clear ensemble-level improvement; the primary gains stem from the dataset expansion and teacher re-fine-tuning.