arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.25058cs.CLcs.CV

ChainDoRA:基于张量列分解的权重分解低秩适配用于参数高效大语言模型微调

ChainDoRA: Tensor-Train Factorized Weight-Decomposed Low-Rank Adaptation for Parameter-Efficient LLM Fine-Tuning

  • Khulna University of Engineering & Technology(库尔纳工程技术大学)
  • Brac University(布拉克大学)
  • Hobart and William Smith Colleges(霍巴特和威廉史密斯学院)

机构由 AI 辅助整理,请以论文原文为准。

Ashfak Yeafi, Mehedi Hasan, Md Khairul Islam

AI总结:

ChainDoRA利用张量列链分解权重方向低秩因子,在LLaMA-7B上以5.35M参数实现72.30%平均准确率,相比DoRA减少90.62%参数并提升推理性能。

AI中文摘要:

参数高效微调(PEFT)在仅更新预训练参数的一小部分的同时,将大语言模型(LLM)适配到下游任务。低秩适配(LoRA)使用两个可训练的低秩矩阵,而权重分解低秩适配(DoRA)进一步分离权重幅度和方向,但在其方向分支中保留了密集的LoRA式分解。我们提出ChainDoRA,一种权重分解适配框架,从连接的张量列(TT)链构建方向性低秩因子,其中适配器秩形成输入侧和输出侧TT收缩之间的边界秩,独立的TT秩控制表示容量和参数成本。在受控的15,119个示例、仅响应适配设置下,使用LLaMA-7B,ChainDoRA与匹配的LoRA和DoRA基线在七个常识推理基准上进行了评估。TT秩为16的ChainDoRA在七个任务上的平均准确率达到72.30%,而LoRA为69.88%,DoRA为69.39%,同时仅需5.35M可训练参数,而LoRA为56.10M,DoRA为56.98M,相对于DoRA减少了90.62%。对TT秩和适配器放置的消融实验显示了可控的参数-准确率权衡,表明连接的TT参数化可以大幅降低幅度-方向适配的参数成本,同时保持并在此设置中提升下游推理性能。

英文摘要:

Parameter-efficient fine-tuning (PEFT) adapts large language models (LLMs) to downstream tasks while updating only a small fraction of their pretrained parameters. Low-Rank Adaptation (LoRA) uses two trainable low-rank matrices, while Weight-Decomposed Low-Rank Adaptation (DoRA) further separates weight magnitude and direction but retains the dense LoRA-style factorization in its directional branch. We propose ChainDoRA, a weight-decomposed adaptation framework that constructs the directional low-rank factors from a connected Tensor-Train (TT) chain, where the adapter rank forms the boundary rank between input- and output-side TT contractions and an independent TT rank controls representation capacity and parameter cost. Under a controlled 15,119-example response-only adaptation setting with LLaMA-7B, ChainDoRA is evaluated against matched LoRA and DoRA baselines on seven commonsense reasoning benchmarks. ChainDoRA with TT rank 16 achieves a seven-task average accuracy of 72.30%, compared with 69.88% for LoRA and 69.39% for DoRA, while requiring only 5.35M trainable parameters versus 56.10M for LoRA and 56.98M for DoRA, corresponding to a 90.62% reduction relative to DoRA. Ablations over TT rank and adapter placement show controllable parameter-accuracy trade-offs, indicating that connected TT parameterization can substantially reduce the parameter cost of magnitude-direction adaptation while preserving, and in this setting improving, downstream reasoning performance.

↑