LoCA:基于局部信用分配的一次性校准后仅前向的大语言模型调优
LoCA: Forward-Only LLM Tuning after One-Shot Calibration with Local Credit Assignment
浏览论文内容
中文总结 AI 辅助
本文提出LoCA方法,通过一次性校准替换大语言模型调优的重复反向传播,在多个基准上优于LoRA,降低了GPU峰值内存、CPU稳态内存与前向传递时间。
中文摘要 AI 辅助
参数高效的后训练可减少可训练参数数量,但仍需通过冻结的骨干网络进行重复的端到端反向传播,因此每次适应步骤都需要具备反向传播能力的硬件,且必须存储或重新计算激活值。本文探讨是否可将这种重复的反向传播链替换为一次性校准,提出了局部信用分配(Local Credit Assignment, LoCA)这一用于小偏移适应的两阶段方法:一次探针反向传播从最终预测误差拟合每个Transformer块的低秩映射,以实现局部隐藏状态校正;LoCA随后复用这些映射,从前向激活值形成分块回归目标,并通过闭式岭回归求解拟合低秩适配器,无需进一步的骨干网络反向传播。在Qwen2.5模型(0.5B至14B)的5个判别基准上评估LoCA,25项任务-规模对比中,16项的评估交叉熵低于对应LoRA运行的结果;包含校准在内的全运行GPU峰值内存比LoRA低26%-29%,校准后CPU稳态内存低36%-52%,每次前向传递时间低43%-48%;共享的规模归一化候选集可在所有测试的Qwen2.5规模及SmolLM2-1.7B上复用,因此LoCA将全局信用分配摊销为一次校准,在重复反向传播不实用的场景下支持后续仅前向调优。本文代码可在此处获取。
英文摘要
Parameter-efficient post-training reduces the number of trainable parameters, but still requires repeated end-to-end backpropagation through the frozen backbone. Every adaptation step therefore needs backward-capable hardware and must store or recompute activations. We ask whether this repeated backward chain can be replaced by a one-time calibration. We introduce Local Credit Assignment (LoCA), a two-stage method for small-shift adaptation. One probe backward pass fits a low-rank map at each transformer block from the final prediction error to a local hidden-state correction. LoCA then reuses these maps to form blockwise regression targets from forward activations and fits low-rank adapters with closed-form ridge solves. No further backbone backward pass is required. We evaluate LoCA on five discriminative benchmarks with Qwen2.5 models from 0.5B to 14B. In 16 of 25 reported task--scale comparisons, LoCA yields lower evaluation cross-entropy than the corresponding LoRA run. Its measured full-run GPU peak, including calibration, is 26--29\% lower than LoRA's. After calibration, its CPU steady-state memory is 36--52\% lower and its per-pass time is 43--48\% lower. A shared scale-normalized candidate set is reused across all tested Qwen2.5 sizes and on SmolLM2-1.7B. LoCA thus amortizes global credit assignment into one calibration and enables later forward-only tuning when repeated backpropagation is impractical. The code associated with this paper is available \href{https://github.com/Xia12121/LoCA}{here}.
发表机构
- University of Oklahoma(俄克拉荷马大学)
- Imperial College London(伦敦帝国学院)
- University of Michigan(密歇根大学)
- Tencent(腾讯)
- University of Edinburgh(爱丁堡大学)
- University of Southern California(南加州大学)
- Beijing Normal University(北京师范大学)
- Beijing Normal-Hong Kong Baptist University(北京师范大学-香港浸会大学联合国际学院)
机构由 AI 辅助整理,请以论文原文为准。