arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

面向资源受限农业田间场景的设备端植物病害检测的轻量级视觉Transformer压缩

Lightweight Vision Transformer Compression for On-Device Plant Disease Detection in Resource-Constrained Agricultural Field Conditions

Mahadev Sunil Kumar, Bhavika Gondi, Desaisetty Venkata Satya Sai Swapnith, Gangireddy Rahul Jogi, Sudheesh Manalil, Arnab Raha, Amitava Mukherjee, Parthasarathy Seethapathy, G. Gopakumar

arXiv 2609.05334首次发表:更新:

发表机构

Amrita Vishwa Vidyapeetham; Intel Corporation; Birla Institute of Technology(阿姆里塔维什瓦维迪亚皮塔姆大学; 英特尔公司; 比拉理工学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对资源受限农业场景的设备端辣椒病害检测,提出结合H-BAC剪枝、量化与注意力知识蒸馏的ViT压缩框架,实现54.5倍体积缩减且保持准确率,明确相关技术的适用场景。

AI 中文摘要

辣椒(Capsicum annuum)是印度经济意义最为重要的作物之一,但其生产力持续受到病害威胁,若无专家介入则难以识别。尽管视觉Transformer(ViTs)已达到很高的分类准确率,但其庞大的计算量使其难以部署在资源受限的设备上。现有压缩方法通常单独处理剪枝、量化和知识蒸馏,未能充分探索它们组合应用的潜在益处与交互作用。我们提出一种统一的视觉Transformer压缩框架,结合基于二阶灵敏度估计的Hessian平衡自适应块剪枝(H-BAC)、量化及基于注意力的知识蒸馏。为系统确定每种压缩类别内的最优配置,我们先通过受控消融研究分别评估各技术,再将表现最佳的组件整合为适配真实农业约束的顺序部署流水线。在包含真实跨村庄、跨设备分布外测试划分的辣椒3类村庄划分数据集上,所得压缩模型的准确率达到或超过95.13%的FP32基线,同时实现74%-98%的模型体积缩减;在四种测试配置下,完全整合的压缩流水线实现54.5倍的体积缩减(从327.42 MB降至6.01 MB),准确率为95.13±2.32%。直接对比进一步显示,在该数据集上,未经过剪枝或蒸馏、直接训练的相同最终规模的学生模型,在相同6.01 MB的INT8体积下达到94.87%的可比准确率,由此明确了H-BAC和知识蒸馏在哪些场景下值得投入计算成本,哪些场景下尚未体现出价值。

英文摘要

Chilli (Capsicum annuum) is one of India's most economically significant crops, yet its productivity is persistently threatened by diseases that are difficult to identify without expert intervention. While Vision Transformers (ViTs) have achieved high classification accuracy, their large computational footprint makes deployment on resource constrained devices challenging. Existing compression approaches typically address pruning, quantization, and knowledge distillation in isolation, leaving the potential benefits and interactions of their combined application insufficiently explored. We propose a unified Vision Transformer compression framework that combines Hessian-Balanced Adaptive Block Pruning (H-BAC), guided by second-order sensitivity estimation, with quantization and attention-based knowledge distillation. To systematically identify the most effective configuration within each compression family, each technique is first evaluated independently through controlled ablation studies, after which the best-performing components are integrated into a sequential deployment pipeline tailored to real-world agricultural constraints. On a chilli 3-class village-split dataset with a genuine cross-village, cross-device out-of-distribution test split, the resulting compressed models match or exceed the 95.13% FP32 baseline's accuracy, alongside 74-98% model size reduction, and the fully integrated compression pipeline achieves a 54.5x size reduction (327.42 MB to 6.01 MB) at 95.13 +/- 2.32% accuracy across four tested configurations. A direct comparison further reveals that, on this dataset, a directly-trained student of the same final size, without pruning or distillation, reaches comparable accuracy of 94.87%, at the same 6.01 MB INT8 size, indicating where H-BAC and knowledge distillation are, and are not yet shown to be, worth their computational cost.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑