发表机构
University of Modena and Reggio Emilia(摩德纳与雷焦艾米利亚大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对设备端植物物种识别的需求,提出超轻量卷积网络BoltNet,结合空间重分配瓶颈与logit预采样,在Pl@ntNet300K等数据集上实现高F1分数,且在多硬件平台上效率最优。
AI 中文摘要
从公民科学图像中自动识别植物物种是一项成熟且要求极高的细粒度识别问题:庞大的分类标签空间、视觉相似的物种以及长尾观测数据,需要模型具备足够的容量,而野外使用场景则对模型的内存、延迟和功耗构成约束。模型规模仅为部署成本的一部分:推理过程中占用内存的中间激活值以及平台相关的执行行为同样重要,因此紧凑的识别模型必须在目标硬件上进行评估,而非仅通过复杂度指标判断。我们提出BoltNet,一种超轻量的全卷积架构,结合空间重分配瓶颈(Spatial Redistribution Bottleneck)和logit预采样(Logit PreSampling),以改善高基数分类中预测性能与模型规模之间的权衡,并提出准确率-压缩权衡(AccuracyCompression Tradeoff)作为补充诊断指标。在Pl@ntNet300K数据集上,BoltNet以34.1万参数(1.37MB)达到0.682的F1分数,是评估的2MB以下模型中最高的F1分数,且接近规模大得多的卷积主干网络。对树莓派5(Raspberry Pi 5)、Jetson Orin Nano和Hailo-8的仅模型测量,表征了其在CPU、GPU和NPU平台上的执行情况,其中BoltNet是始终效率最高的模型,在GPU和NPU上具有最佳的FPS/W,在CPU上为次佳。在AIDERv2和CLRS上的结果提供了其在环境图像分类任务中可迁移的辅助证据。代码可访问:this https URL
英文摘要
Automated plant species identification from citizen-science imagery is an established, demanding fine-grained recognition problem: large taxonomic label spaces, visually similar species, and long-tailed observations require real model capacity, while field use constrains memory, latency, and power. Model size is only part of the deployment cost: intermediate activations held in memory during inference and platformdependent execution behavior matter too, so compact recognition must be assessed on target hardware rather than through complexity metrics alone. We present BoltNet, an ultra-lightweight fully convolutional architecture combining a Spatial Redistribution Bottleneck and Logit PreSampling to improve the tradeoff between predictive performance and model size in high-cardinality classification, and report the AccuracyCompression Tradeoff as a complementary diagnostic. On Pl@ntNet300K, BoltNet reaches 0.682 F1-score with 341K parameters (1.37 MB), the highest F1-score among evaluated models below 2 MB and close to substantially larger convolutional backbones. Model-only measurements on a Raspberry Pi 5, Jetson Orin Nano, and Hailo-8 characterize execution across CPU, GPU, and NPU platforms, where BoltNet is the most consistently efficient model, with the best FPS/W on the GPU and NPU and second-best on the CPU. Results on AIDERv2 and CLRS provide secondary evidence of transfer across environmental image-classification tasks. Code available at: https://codeberg.org/danielrossi/BoltNet
CommentsAccepted at the CVPPA (Computer Vision Problems in Plant Phenotyping and Agriculture) workshop, ECCV 2026. 17 pages, 2 figures