发表机构
Utrecht University; University of Twente(乌得勒支大学; 特文特大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究通过评估25个开源LLM在RepoBench和McEval上的表现,揭示了代码补全中能耗主要由任务结构决定,较小且量化模型常实现帕累托最优,为可持续AI辅助开发提供了节能部署策略。
AI 中文摘要
代码补全是大型语言模型(LLM)在软件开发中应用最广泛的场景之一。由于隐私方面的考虑,开源权重LLM越来越多地被用于本地部署的代码补全系统。尽管LLM的准确性有所提升,但推理的能耗成本仍未得到充分探索,尤其是在大上下文工作负载和跨编程语言场景下。本研究探讨了基于LLM的代码补全中准确性与能耗之间的权衡,以及工作负载特征、上下文大小和模型规模如何影响推理能耗。我们在两个工作负载上评估了25个开源权重LLM:在RepoBench上,针对仓库级下一行补全,采用不同上下文大小;在McEval上,针对Python、Java和Rust的填充中间(FIM)代码补全。我们使用相关性分析和聚类稳健线性回归,分析了输入令牌、输出令牌、模型大小及其交互对能耗的影响。我们的研究结果表明,能耗的主要驱动因素强烈依赖于任务结构。在RepoBench中,能耗主要受输入上下文大小及其与模型规模的交互影响,而在McEval中,输出生成及其与激活参数数量的交互占主导地位。输出生成每令牌的能耗高于提示处理。在两个基准测试中,较小且高度量化的模型经常实现帕累托最优的权衡,通常提供与较大FP16模型相当的准确性,同时消耗显著更少的能量。增加模型大小或上下文长度并不一定带来成比例更好的补全质量,而量化可以在有限的准确性下降下大幅提高能源效率。这些发现支持更节能的部署策略,以促进可持续的AI辅助软件开发。
英文摘要
Code completion is one of the most widely used applications of large language models (LLMs) in software development. Open-weight LLMs are increasingly adopted for locally deployed code completion systems, partly due to privacy concerns. Despite advances in LLM accuracy, the energy cost of inference remains underexplored, particularly under large-context workloads and across programming languages. This study investigates the trade-off between accuracy and energy consumption in LLM-based code completion and how workload characteristics, context size, and model scale influence inference energy usage. We evaluate 25 open-weight LLMs on two workloads: repository-level next-line completion with varying context sizes on RepoBench, and fill-in-the-middle (FIM) code completion across Python, Java, and Rust on McEval. We analyze the influence of input tokens, output tokens, model size, and their interactions on energy consumption using correlation analysis and cluster-robust linear regression. Our findings show that the dominant drivers of energy consumption depend strongly on task structure. In RepoBench, energy consumption is primarily influenced by input context size and its interaction with model scale, whereas in McEval, output generation and its interaction with active parameter count dominate. Output generation is more energy-intensive per token than prompt processing. Across both benchmarks, smaller and heavily quantized models frequently achieve Pareto-optimal trade-offs, often providing accuracy comparable to larger FP16 models while consuming substantially less energy. Increasing model size or context length does not necessarily lead to proportionally better completion quality, while quantization can substantially improve energy efficiency with limited accuracy degradation. These findings support more energy-aware deployment strategies for sustainable AI-assisted software development.