TASTE:面向设备端边缘学习的吞吐量感知批量大小调优
TASTE: Throughput-Aware Batch Size Tuning for On-Device Edge Learning
- FZI Research Center for Information Technology(FZI信息技术研究中心)
- University of Tübingen(蒂宾根大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文提出TASTE方法,利用贝叶斯优化调整批量大小以最大化边缘设备训练吞吐量,实验表明最佳批量大小配合梯度累积和线性学习率缩放可提升高达2倍吞吐量,并在持续学习中保持稳定性-可塑性平衡。
AI中文摘要:
隐私保护人工智能(AI)的兴起已将模型适配和个性化的焦点转向设备端学习,即在边缘硬件上直接使用本地用户数据对深度学习模型进行微调。然而,这种转变需要在资源受限的硬件上优化深度学习训练,以在保持预测准确性的同时最大化吞吐量。本文介绍了一种用于设备端模型训练的新技术,该技术采用一种高效的基于贝叶斯优化的批量大小调优方法,以最大化硬件吞吐量。为了评估该超参数对学习动态的影响,我们研究了两种不同的范式:标准监督学习(SL)和在线持续学习(CL)。在多种边缘设备上的实验结果表明存在一个吞吐量上限,超过该上限后,增加批量大小不会带来额外的吞吐量提升。所提出的调优方法能够确定最佳批量大小,当与梯度累积和线性学习率缩放相结合时,在Raspberry Pi 4等平台上,与最大批量大小相比,训练吞吐量可提升高达2倍,且不损害模型准确性。此外,在CL范式中,我们证明了最佳批量大小能够维持增量学习所需的稳定性-可塑性平衡,在最大化边缘硬件计算效率的同时有效缓解灾难性遗忘。
英文摘要:
The rise of privacy-preserving artificial intelligence (AI) has shifted the focus of model adaptation and personalization towards on-device learning, where deep learning models are finetuned directly on edge hardware using local user data. However, this shift requires optimization of deep learning training on resource-constrained hardware to maximize throughput while maintaining predictive accuracy. This paper introduces a novel technique for on-device model training that incorporates an efficient Bayesian optimization-based batch size tuning approach to maximize hardware throughput. To evaluate the impact of this hyperparameter on the learning dynamics, we investigated two distinct paradigms: standard supervised learning (SL) and online continual learning (CL). Experimental results across various edge devices demonstrate a throughput ceiling, beyond which increasing the batch size yields no additional throughput gains. The proposed tuning approach identifies the optimal batch size, which, when combined with gradient accumulation and linear learning rate scaling, achieves up to a 2X increase in training throughput on platforms such as Raspberry Pi 4 compared to maximum batch sizes, without compromising model accuracy. Furthermore, in the CL paradigm, we demonstrate that optimal batch sizes maintain the stability-plasticity balance required for incremental learning, effectively mitigating catastrophic forgetting while maximizing computational efficiency on edge-hardware.