发表机构
AURA Lab Department of Computer Science William \& Mary Williamsburg, VA, USA
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究在资源受限硬件上运行大型代码模型时,六种量化方法对代码质量的影响,用多基准评估功能正确性等指标,引入提示复杂性分析,揭示不同量化技术在正确性、代码质量及鲁棒性上的差异,为选择量化策略提供指导。
AI 中文摘要
随着Ollama等本地推理框架的广泛应用,开发者常在笔记本电脑等资源受限硬件上运行大型代码模型。训练后量化对减少内存占用和实现实际部署至关重要,但对生成代码的影响尚不清楚。本文在两个大型代码模型家族Qwen2.5-Coder和CodeLlama上,用多语言McEval和CoderEval基准对六种量化方法进行实证评估,评估功能正确性及可维护性、可靠性、安全性和结构复杂性等。还引入新分析,结果表明量化技术对正确性和代码质量影响不同,AQLM表现出色,QuIP#正确性下降大,安全属性稳定,对提示复杂性的鲁棒性因技术而异。
英文摘要
The growing adoption of local inference frameworks such as Ollama has made it increasingly common for developers to run large code models on laptops and other resource-constrained hardware. In these settings, post-training quantization is essential for reducing memory footprint and enabling practical deployment, yet its impact on generated code remains insufficiently understood. We empirically evaluate six state-of-the-art quantization methods (GPTQ, AWQ, QuIP#, AQLM, BitsAndBytes, and GGUF) on two representative large code model families, Qwen2.5-Coder and CodeLlama, using the multilingual McEval and CoderEval benchmarks for Python and Java. We assess functional correctness (pass@1) together with maintainability, reliability, security, and structural complexity. We also introduce a novel analysis of robustness under varying prompt complexity, characterized by Shannon entropy and token length. Our results show that quantization techniques differ meaningfully in their impact on correctness and code quality. AQLM consistently matches or exceeds the full-precision baseline, whereas QuIP# exhibits the largest correctness degradation, particularly on complex prompts. Security attributes remain stable across models, benchmarks, and programming languages, while robustness to prompt complexity varies across techniques. These findings provide practical guidance for selecting quantization strategies for deploying large code models on resource-constrained hardware and highlight the importance of evaluating quantized models beyond functional correctness.