arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

鸟类物种识别需要何种输入分辨率,其在边缘设备上的延迟成本是多少?一项包含14种输入分辨率和六种架构的实测研究

What Input Resolution Is Required for Bird Species Identification, and What Is Its Latency Cost on an Edge Device? A Study of 14 Input Resolutions and Six Architectures with On-Device Measurements

Takeshi Nishikawa

arXiv 2609.14247首次发表:更新:

发表机构

Foundation for Computational Science (FOCUS)(计算科学基金会(FOCUS))

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究通过14种输入分辨率、六种架构及边缘设备实测,发现目标精度决定最优分辨率,更换模型比提高分辨率更有效,并揭示了预处理路径、精度引擎及FP16数值范围对延迟和精度的重要影响。

AI 中文摘要

风电场鸟类撞击缓解需要识别仅占几十像素的远处鸟类,因此分类器的输入分辨率N是一个设计变量,而非固定规格。我们通过一个因子设计对其进行研究,涵盖14种边长N(16至224)、六种架构、两种训练与评估机制以及30个随机种子——共2,520个检查点和5,040次评估——并在NVIDIA Jetson Orin Nano上测量延迟。得到四项结果。(1) 所选N取决于目标:ResNet50在N=112时于验证集上达到0.90,估计耗时1.85毫秒(测试集上为0.8980);DINOv2-L在N=144时达到0.95,耗时12.70毫秒;更换模型比提高N能带来更多精度提升(在N=112时,+5.93对比+2.33个百分点)。(2) 降低N的收益取决于假定的预处理路径:当每个个体从其自身文件解码时,N从224降至80节省13.5%的时间;但当检测器一次性解码4K帧时,节省46.2%的时间;帕累托最优集从20个配置增加到23个。(3) 精度必须在部署引擎上测量:半精度单独使ViT-S/16在N>=96时损失4至7个百分点,而CNN的损失保持在0.1个百分点以内;在验证时保持选择的情况下,176个目标中有26个的选择不同。一个损坏的FP16引擎可能比正确的引擎运行得更快,这从延迟上无法察觉;允许14个ViT-S/16 FP32配置将推荐结果移至0.931-0.938区间内,并满足10毫秒预算。(4) ViT-L规模的模型适合此设备,但激活值超出FP16范围;在Transformer块边界处分割计算图,将FP32限制在受影响的段内,使部署的DINOv2-L链比单引擎构建快1.85倍。我们还量化了机制差异符号如何随种子数量而稳定;一种敏感性分割移除了某些形式的组共享,在选择边界处保留了所有14个非平凡符号。

英文摘要

Bird-strike mitigation at wind farms requires identifying distant birds that span only tens of pixels, so the classifier's input resolution N is a design variable, not a fixed specification. We study it with a factorial design over 14 side lengths N (16 to 224), six architectures, two training and evaluation regimes and 30 random seeds -- 2,520 checkpoints and 5,040 evaluations -- plus latency measured on an NVIDIA Jetson Orin Nano. Four results. (1) The selected N depends on the target: 0.90 is met on validation by ResNet50 at N=112 in an estimated 1.85 ms (0.8980 on test) and 0.95 by DINOv2-L at N=144 in 12.70 ms; changing the model buys more accuracy than raising N (+5.93 versus +2.33 points at N=112). (2) The benefit of lowering N depends on the assumed preprocessing path: N=224 -> 80 saves 13.5% when each individual is decoded from its own file but 46.2% when the detector decodes the 4K frame once; the Pareto set grows from 20 to 23 configurations. (3) Accuracy must be measured on the deployed engine: half precision costs ViT-S/16 alone 4 to 7 points at N>=96 while the CNNs stay within 0.1 points, and with selection held at validation the choice differs at 26 of 176 targets. A broken FP16 engine can run faster than a correct one, undetectable from latency; admitting 14 ViT-S/16 FP32 configurations moves the recommendation over the 0.931-0.938 band and under the 10 ms budget. (4) ViT-L-scale models fit this device, but activations exceed the FP16 range; splitting the graph at transformer-block boundaries confines FP32 to the affected segments, making the deployed DINOv2-L chain 1.85x faster than the single-engine build. We also quantify how the regime-difference sign stabilises with seed count; a sensitivity split removing some forms of group sharing preserves all 14 non-trivial signs at the selection boundary.

Comments25 pages, 7 figures, 7 tables. On-device measurements on NVIDIA Jetson Orin Nano (JetPack 6.2, TensorRT 10.3). Data and code: https://doi.org/10.5281/zenodo.22697995

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑