arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

面向茶叶病害分类的、从视觉基础模型到轻量视觉状态空间模型的跨架构知识蒸馏

Cross-Architecture Knowledge Distillation from a Vision Foundation Model to a Lightweight Visual State Space Model for Tea Leaf Disease Classification

Zibo Zhou, Zongsen Qiu, Rui Chen, Yujie Yao, Yue Zhou, Jianjun Wang

arXiv 2608.26771首次发表:更新:

AI 中文总结

本文针对茶叶病害分类,通过跨架构知识蒸馏将DINOv2教师模型的知识迁移到轻量LVSSM学生模型,提升了边缘部署的轻量模型准确率,参数仅为教师模型的1/5,准确率保留98.3%。

AI 中文摘要

茶叶病害自动分类支持精准农业,但在有限计算资源下将高精度模型部署到边缘设备仍具挑战性。DINOv2等自监督视觉基础模型能提供强特征,但体积过大难以在田间部署;而在小型农业数据集上从头训练的轻量模型往往存在欠拟合问题。本文研究跨架构知识蒸馏(KD),即从微调后的DINOv2教师模型(视觉Transformer)向紧凑的双向视觉状态空间模型(LVSSM)学生模型进行知识迁移,该方向研究较少,因两类架构采用完全不同的令牌混合机制。我们识别并修复了两个导致从头训练的SSM学生模型在有限数据上无法学习的训练稳定性问题:单个大尺寸补丁嵌入卷积,以及破坏残差路径的融合层。通过采用渐进式卷积主干和门控双向选择性扫描块,参数规模为445万的学生模型可稳定训练。在三个随机种子下,经温度缩放的logit蒸馏使测试准确率从92.32±2.14%提升至95.41±1.17%(最优单轮准确率为96.20%,宏F1值为94.45%),平均提升3.09个百分点。该学生模型参数仅为2200万参数教师模型的1/5,却保留了教师模型98.3%的准确率。消融实验显示,中间特征对齐损失会降低准确率,因此简单的logit级KD是最优配置。公平的从头训练对比表明,该增益仅适用于初始性能低于教师模型的学生。我们报告了每类指标、混淆矩阵、自助法置信区间及FLOPs/延迟测量值,并讨论了局限性,包括单数据集范围和简化的非官方SSM实现。

英文摘要

Automated tea leaf disease classification supports precision agriculture, yet deploying accurate models on edge devices remains challenging under tight compute budgets. Self-supervised vision foundation models such as DINOv2 provide strong features but are too large for field deployment, while lightweight models trained from scratch on small agricultural datasets often underfit. We study cross-architecture knowledge distillation (KD) from a fine-tuned DINOv2 teacher (Vision Transformer) to a compact bidirectional Visual State Space Model (LVSSM) student, an underexplored direction because the architectures use fundamentally different token-mixing mechanisms. We identify and fix two training-stability problems that prevent the from-scratch SSM student from learning on limited data: a single large patch-embedding convolution and a fusion layer that severs the residual path. With a progressive convolutional stem and gated bidirectional selective-scan block, the 4.45M-parameter student trains stably. Across three seeds, temperature-scaled logit distillation raises test accuracy from 92.32+/-2.14% to 95.41+/-1.17% (best single run: 96.20%; macro-F1: 94.45%), a +3.09 percentage-point mean gain. The student uses 5.0 times fewer parameters than the 22M-parameter teacher while retaining 98.3% of its accuracy. Ablations show that intermediate feature-alignment losses reduce accuracy, making simple logit-level KD the strongest configuration. A fair from-scratch comparison shows the gain is specific to students that start below the teacher. We report per-class metrics, confusion matrices, bootstrap confidence intervals, and FLOPs/latency measurements, and discuss limitations including the single-dataset scope and simplified non-official SSM implementation.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑