arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

你能做多小?关于60M参数模型上文本到SQL的LoRA秩、目标模块和量化权衡的对照研究

How Small Can You Go? A Controlled Study of LoRA Rank, Target Modules, and Quantization Trade-offs for Text-to-SQL on a 60M-Parameter Model

Mahendra Singh Rathor, Anagheem Azzam

arXiv 2607.25583首次发表:更新:

AI 中文总结

研究在60M参数模型上文本到SQL时,LoRA秩、目标模块和量化的权衡。通过对照单变量研究,发现r = 16的LoRA在训练参数少、内存消耗低时接近全量微调准确性,QLoRA的INT8和NF4量化以低内存成本实现可比准确性,为内存受限部署提供方案。

AI 中文摘要

参数高效微调(PEFT)和低比特量化是在计算预算紧张时适配语言模型的标准工具,但其相互作用大多在十亿参数模型上研究,设计空间探索成本高。本文提出互补问题:在特定的、完全可重现的60M参数编解码器模型(T5-small)和单表文本到SQL基准(WikiSQL)上,每个效率旋钮实际会牺牲多少任务准确性?通过对LoRA秩r(取值为2、4、8、16、32)、适配模块集和数值精度进行对照单变量研究,报告任务准确性及系统级指标。结果表明,r = 16的LoRA在训练参数少于1%且峰值GPU内存消耗减少31%的情况下,恢复到全量微调准确性的11.6个百分点内。在此设置下,r超过16无显著准确性提升。QLoRA的INT8和NF4量化以显著更低内存成本实现可比准确性,为内存受限部署提供了有吸引力的权衡。所有代码、配置和日志均已发布以实现完全可重现性。

英文摘要

Parameter-efficient fine-tuning (PEFT) and low-bit quantization are now standard tools for adapting language models under tight compute budgets, yet their interaction is most often studied on billion-parameter models where the design space is expensive to explore. We ask a complementary question: on a specific, fully reproducible 60M-parameter encoder-decoder model (T5-small) and a single-table text-to-SQL benchmark (WikiSQL), how much task accuracy does each efficiency knob actually cost? We run a controlled, single-variable study over (i) LoRA rank r in {2, 4, 8, 16, 32}, (ii) the set of adapted modules, and (iii) numerical precision. We report task accuracy alongside system-level metrics including trainable parameters, peak training memory, inference latency, and throughput, and frame adaptation as a constrained trade-off rather than an accuracy-only objective. Our results show that LoRA with r=16 recovers within 11.6 percentage points of full fine-tuning accuracy (59.6% vs. 71.2% exact-match) while training fewer than 1% of parameters and consuming 31% less peak GPU memory. Within this setting, rank beyond r=16 yields no measurable accuracy gain. QLoRA with INT8 and NF4 quantization achieves comparable accuracy (52.8% and 53.2%) at dramatically lower memory cost (0.60 GB each), demonstrating a compelling trade-off for memory-constrained deployments. All code, configurations, and logs are released for full reproducibility.

Comments8 pages, 3 figures, 3 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑