arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SPRKD:通过鞍点近似实现深度神经网络的有效知识蒸馏

SPRKD: Effective Knowledge Distillation for Deep Neural Networks via Saddle Region Approximation

Aditya Dewan, Arjun Yogeswaran, Benjamin Fedoruk

arXiv 2607.23346首次发表:更新:

AI 中文总结

研究针对深度神经网络参数多阻碍低计算环境部署及现有知识蒸馏方法不足的问题;提出SPRKD方法,利用鞍点近似,通过黑塞ESD识别鞍区域等技术;在多数据集上取得良好效果,收敛特性优,提升了知识蒸馏性能。

AI 中文摘要

现代深度神经网络对科学和工业影响巨大,但过多参数阻碍其在低计算环境中的部署。主要的知识蒸馏(KD)方法倾向于复制,导致学生模型任务性能低、收敛受阻。我们提出鞍点招募知识蒸馏(SPRKD),将蒸馏从复制转变为利用教师模型作为优化曲率和域代理,通过嵌入和盆地分形特性将鞍点表征为具有强大进一步下降潜力的区域。使用黑塞特征值谱密度(ESD),SPRKD识别低损失鞍区域供学生重新探索;弱教师集合聚合成近似鞍区域(ASR),通过注入式迁移学习重新参数化到学生模型,并采用指数衰减欧几里得变换、负黑塞特征步长和高斯扰动逼近。在从弱的25546参数教师模型蒸馏得到的6430参数CNN对疟疾血涂片分类任务中,SPRKD达到94.8%的验证准确率,优于响应KD 24.70个百分点,与相同架构的从头训练基线在统计上等效。在MNIST、CIFAR - 100和TinyImageNet上,SPRKD在初步基准测试中比从头训练基线高出多达8个百分点。黑塞ESD和二维损失景观分析表明,与响应KD和对照学生模型相比,SPRKD收敛到更宽的最小值,黑塞迹和谱半径显著更小,表明下降更平滑且噪声鲁棒性更强。

英文摘要

Modern deep neural networks are potent catalysts for scientific and industrial impact, yet excessive parameter counts impede deployment in low-compute settings such as hospital equipment and energy infrastructure. Predominant knowledge distillation (KD) methods favor replication: smaller students mimic teacher output logits, yet empirically yield low task performance, hamper convergence, and act merely as regularization rather than substantive knowledge transfer. We propose Saddle Point Recruitment for Knowledge Distillation (SPRKD), reframing distillation from replication to employing teachers as optimization-curvature and domain proxies, characterizing saddle points as regions of strong further-descent potential via embedding and basin-fractal properties. Using Hessian eigenvalue spectral density (ESD), SPRKD identifies low-loss saddle regions for student re-exploration; weak-teacher ensembles are aggregated into an Approximated Saddle Region (ASR), re-parameterized into the student via Transfer Learning by Injection, and approached with exponentially decaying Euclidean transformations, Negative Hessian Eigensteps, and Gaussian perturbations. On malaria blood smear classification with a 6,430-parameter CNN distilled from a weak 25,546-parameter teacher, SPRKD reaches 94.8% validation accuracy, outperforming Response KD by 24.70 percentage points (McNemar p = 6.3e-87) and matching scratch-trained baselines of the same architecture to statistical equivalence (p = 1.0). Across MNIST, CIFAR-100, and TinyImageNet, SPRKD exceeds scratch-trained baselines by up to 8 percentage points on preliminary benchmarks. Hessian ESD and 2-D loss landscape analysis show convergence to wider minima with substantially smaller Hessian trace and spectral radius than Response KD and control students, indicating smoother descent and greater noise robustness.

Comments12 pages, 5 figures. Code and reproduction scripts: https://github.com/thetechdude124/SADDLE-POINT-RECRUITMENT-FOR-KNOWLEDGE-DISTILLATION. Awarded 2nd Place in Mathematical and Cybersecurity Research (NSA) at Regeneron ISEF 2023 and an Outstanding Research Award at WAICY (World AI Competition for Youth) 2022

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑