arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

面向消费级GPU上个性化端侧小型语言模型(SLM)的高能效、低内存参数高效微调(PEFT)方法

Energy- and Memory-Efficient PEFT Methods for Personalized On-Device SLMs on Consumer GPUs

Kuanysh Akhmetzhanov, Jurn-Gyu Park

arXiv 2608.04488首次发表:更新:

发表机构

Nazarbayev University(纳扎尔巴耶夫大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对消费级GPU上的个性化端侧SLM,对比五种PEFT方法在四类SLM、六项基准任务的能耗与内存表现,发现LoRA+是多数场景的最优PEFT方法,QLoRA在内存受限场景具优势,为端侧个性化LLM部署提供了能耗与内存平衡的可行方案。

AI 中文摘要

尽管大型语言模型(LLM)发展迅速,但由于显存(VRAM)、时间和能耗成本高昂,在资源受限设备上部署和个性化定制LLM仍不现实。小型语言模型(SLM)的参数高效微调(PEFT)提供了一种有前景的替代方案,但很少有研究在考虑能耗的情况下,针对不同架构、同时使用通用基准和个性化基准对PEFT方法进行比较。我们在两类共四款SLM上比较了五种微调方法(全量微调、LoRA、LoRA+、QLoRA和BitFit),这两类SLM分别是基于Transformer的TinyLlama-1.1B、Qwen3-1.7B,以及基于状态空间模型(SSM)的Mamba-1.4B、Mamba-2-1.3B;评估任务涵盖三个通用语言理解评估(GLUE)任务(SST-2、QNLI、STS-B)和三个LaMP个性化任务(LaMP-1、LaMP-2、LaMP-3)。每个配置均采用以能耗为核心的NetScore-E和以内存为核心的NetScore-M进行评估,这两个指标反映了端侧部署的核心约束。我们采用严格的能耗优先规则选择方法(优先选择NetScore-E最高的,若存在平局则由NetScore#决定)。LoRA+在24种配置中19种的NetScore-E最高,13种的NetScore-M最高,且在24种配置中18种被选为最优方法。仅适用于Transformer模型的QLoRA,其微调峰值VRAM相比LoRA最多降低3.9倍,因此在12种Transformer配置中5种的NetScore-M最优,但由于反量化开销,在能耗优先的选择规则下仅1种配置选用QLoRA。BitFit和全量微调在两个指标上几乎没有竞争力,而TinyLlama-1.1B在6个基准中的5个能耗导向NetScore-E最高,在4个内存导向NetScore-M最高。这些结果表明,紧凑SLM结合PEFT为个性化端侧部署提供了一条实用的能耗感知路径,最优方法由主导约束决定:能耗优先选LoRA+,内存优先选QLoRA。

英文摘要

Despite rapid advances in large language models (LLMs), deploying and personalizing them on resource-constrained devices remains impractical due to high VRAM, time, and energy costs. Parameter-Efficient Fine-Tuning (PEFT) of Small Language Models (SLMs) offers a promising alternative, yet few studies compare PEFT methods across architectures using both general and personalization benchmarks while accounting for energy consumption. We compare five fine-tuning approaches (Full Fine-Tuning, LoRA, LoRA+, QLoRA, and BitFit) on four SLMs from two families (Transformer-based: TinyLlama-1.1B, Qwen3-1.7B; SSM-based: Mamba-1.4B, Mamba-2-1.3B) across three GLUE tasks (SST-2, QNLI, STS-B) and three LaMP personalization tasks (LaMP-1, LaMP-2, LaMP-3). Each configuration is evaluated with the energy-focused NetScore-E and the memory-focused NetScore-M, the two variants that reflect the constraints binding on-device deployment. Methods are selected with a strict energy-first rule (highest NetScore-E, ties broken by NetScore#). LoRA+ achieves the highest NetScore-E in 19 of 24 configurations and the highest NetScore-M in 13 of 24, and is the selected method in 18 of 24. QLoRA, available only for the Transformer models, cuts peak finetuning VRAM by up to 3.9x relative to LoRA and therefore takes the best NetScore-M in 5 of the 12 Transformer configurations, although its de-quantization overhead leaves it selected in only one of them once energy decides. BitFit and full fine-tuning are almost never competitive on either variant, and TinyLlama-1.1B leads the energy-focused NetScore-E on five of the six benchmarks and the memory-focused NetScore-M on four. These results show that compact SLMs paired with PEFT provide a practical, energy-aware path to personalized on-device deployment, with the optimal method set by the dominant constraint: LoRA+ for energy and QLoRA for memory.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑