arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.33548cs.LGcs.AIeess.SP

TerMeZO:面向边缘BitNet模型微调的三元稀疏零阶优化

TerMeZO: Ternary Sparse Zeroth-Order Optimization for Fine-tuning BitNet Models at the Edge

Houssem Sifaou, Prabodh Katti, Bipin Rajendran, Osvaldo Simeone

首次发表
浏览论文内容

中文总结 AI 辅助

针对BitNet模型微调内存开销大的问题,提出TerMeZO,利用三元量化器几何特性构建稀疏掩码,无需额外数据或内存,实现更快收敛并匹配或超越全参数MeZO性能。

中文摘要 AI 辅助

使用一阶优化器微调大语言模型(LLMs)所需的内存比推理所需的内存大数倍。内存高效的零阶优化(MeZO)仅通过前向传播估计梯度,从而绕过了这一成本。然而,对于BitNet架构(一类具有三元{-1,0,1}权重和8位激活的LLM),微调需要更新全精度潜在权重,因此MeZO的内存占用不再与推理相匹配。一个有前景的解决方案是仅微调潜在权重的子集,但现有的稀疏零阶(ZO)方法要么忽略三元结构,要么需要一阶梯度信息来构建稀疏掩码,这与ZO微调的目的相悖。我们提出了TerMeZO,一种稀疏MeZO方案,它利用三元量化器本身的几何特性来识别在微调期间更可能改变值的潜在权重,且无需额外的数据或内存成本。我们的收敛性分析表明,由于TerMeZO优化了微调有效维度的缩减,其收敛速度可以快于全参数MeZO。我们在参数规模从1B到3B的BitNet模型上进行了大量实验,涵盖分类、指令遵循和数学推理任务。TerMeZO在显著降低微调内存占用的情况下,匹配或超过了全参数MeZO的性能。

英文摘要

Fine-tuning anguage models (LLMs) with first-order optimizers requires a memory several times larger than that required for inference. Memory-efficient zeroth-order optimization (MeZO) sidesteps this cost by estimating gradients from forward passes only. However, for BitNet architectures, a family of LLMs with ternary {-1,0,1\} weights and 8-bit activations, fine-tuning requires updating full-precision latent weights, and thus the memory footprint of MeZO no longer matches that of inference. A promising solution is to finetune only a subset of the latent weights, but existing sparse zeroth-order (ZO) methods either ignore the ternary structure or require first-order gradient information to build a sparse mask, which is at odds with the purpose of ZO fine-tuning. We propose TerMeZO, a sparse MeZO scheme that exploits the geometry of the ternary quantizer itself to identify the latent weights that are more likely to change values during fine-tuning, at no additional data or memory cost. Our convergence analysis shows that TerMeZO can converge faster than full-parameter MeZO, owing to its optimized reduction of the fine-tuning effective dimension. We run extensive experiments on BitNet models ranging from 1B to 3B parameters, spanning classification, instruction-following, and mathematical reasoning tasks. TerMeZO matches or exceeds the performance of full-parameter MeZO while substantially reducing the fine-tuning memory footprint.

发表机构

  • Northeastern University London(伦敦东北大学)

机构由 AI 辅助整理,请以论文原文为准。

↑