arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

QScheduler:用于在 INT8 NPU 上进行零阶设备端训练的自适应梯度采样

QScheduler: Adaptive Gradient Sampling for Zeroth-Order On-Device Training on INT8 NPUs

Victor Felipe Domingues Do Amaral, Pierre Demaj, Erwan Libessart, Laurent Folliot, Anthony Kolar, Philippe Bénabès

arXiv 2607.18802首次发表:更新:

AI 中文总结

研究在INT8 NPU上零阶设备端训练中梯度样本数量q的影响及优化难题,提出自适应算法QScheduler,可依训练进度调整q,实验表明该算法在EuroSAT和STL-10数据集上,与固定q配置效果相当且无需事先优化超参数。

AI 中文摘要

零阶(ZO)优化通过仅前向传递估计梯度,实现了在配备 NPU 的微控制器上进行设备端学习(ODL),无需反向传播原语并降低内存需求。梯度样本数量q对训练有关键影响:样本不足会产生噪声梯度导致早期停滞,样本过多则消耗更多计算资源。然而,找到最优q通常需要昂贵的超参数搜索。本文介绍了QScheduler,一种基于训练进度调整q的自适应算法,并首次在STM32N6的Neural-ART NPU上提供了INT8量化设备端训练的概念验证。在EuroSAT和STL-10上的实验表明,QScheduler在ResNet18和MobileNetV2上与调优良好的固定q配置匹配良好,无需事先进行q超参数优化。

英文摘要

Zeroth-Order (ZO) optimization enables On-Device Learning (ODL) on NPU-equipped microcontrollers by estimating gradients through forward passes alone, bypassing the need for backpropagation primitives and reducing memory requirements. The number of gradient samples q critically affects training: insufficient samples produce noisy gradients that plateau early, while excessive samples consume more computational resources. However, finding an optimal q typically requires costly hyperparameter searches. This work introduces QScheduler, an adaptive algorithm that adjusts q based on training progress, and provides the first proof-of-concept of INT8 quantized on-device training on the STM32N6's Neural-ART NPU. Experiments on EuroSAT and STL-10 show that QScheduler matches well-tuned fixed-q configurations for both ResNet18 and MobileNetV2, without requiring prior q hyperparameter optimization.

Journal refInternational Joint Conference on Neural Networks, Jun 2026, Maastricht, Netherlands

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑