arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

超越FLOPs:面向代码相关任务的可持续大语言模型的能量感知知识蒸馏

Beyond FLOPs: Energy-Aware Knowledge Distillation for Sustainable LLMs on Code-Related Task

Enrique Barba Roque, Luís Cruz, Annibale Panichella

arXiv 2608.17515首次发表:更新:

发表机构

Delft University of Technology(代尔夫特理工大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对代码相关任务,发现FLOPs并非可靠的能耗指标,提出结合能量代理模型的Morph知识蒸馏方法,可将学生模型推理能耗降90%、内存降86%,实现LLMs的可持续部署。

AI 中文摘要

背景:大语言模型(LLMs)正越来越多地应用于软件工程(SE)任务,在克隆检测、漏洞预测和代码摘要生成等问题上取得了较高的准确率。然而,它们的高计算需求和能耗引发了可持续性方面的担忧,并阻碍了其在消费级硬件和资源受限平台上的应用。在文献和工业界,报告LLMs计算成本的常用方式是使用模型前向传播所需的浮点运算次数(FLOPs)。目标:本文研究能量感知知识蒸馏在软件工程中的应用意义,旨在在保持性能的同时提高模型效率,并确定FLOPs是否是可靠的能量感知度量指标。方法:我们使用基于多目标优化的蒸馏方法Morph开展对照实验,以实证检验FLOPs在克隆检测和漏洞预测任务中是否能准确反映能耗。我们将该方法扩展为包含能量代理模型,这些模型可在优化过程中直接估计中央处理器(CPU)和图形处理器(GPU)的能耗,并将Morph应用于使用CodeT5+的代码摘要生成这类生成式任务。结果:我们的结果表明,FLOPs并非始终是能耗的可靠指标,而使用能量代理模型可取得更好的结果。蒸馏得到的学生模型可将推理能耗降低多达90%,内存使用量降低86%,仅伴随适度的准确率权衡。结论:由直接能量代理而非FLOPs引导的能量感知知识蒸馏,可改善软件工程应用中LLMs的能耗、可持续性和可部署性,使消费级硬件上能够运行高效模型。

英文摘要

Background: Large Language Models (LLMs) are increasingly being applied to Software Engineering (SE) tasks, achieving high accuracy across problems such as clone detection, vulnerability prediction, and code summarization. However, their high computational demands and energy consumption raise sustainability concerns and hinder their use on consumer hardware and resource-constrained platforms. A common way to report the computational cost of an LLM in the literature and industry is to use the number of Floating Point Operations (FLOPs) required to perform a pass over the network. Aims: This paper investigates the implications of energy-aware knowledge distillation for SE, aiming to improve model efficiency while maintaining performance and to determine whether FLOPs is a reliable energy-aware metric. Method: We conduct a controlled experiment using Morph, a Many-Objective Optimization-based distillation methodology, to empirically examine whether FLOPs accurately reflect energy consumption in Clone Detection and Vulnerability Prediction tasks. We extend this methodology to include energy-surrogate models that directly estimate CPU and GPU energy consumption during optimization, and we apply Morph to generative tasks using CodeT5+ for code summarization. Results: Our results show that FLOPs is not always a reliable indicator of energy consumption, and better results can be achieved by using energy-surrogate models. Distilled student models can reduce inference energy consumption by up to 90\% and memory usage by 86\%, with only modest accuracy trade-offs. Conclusions: Energy-aware knowledge distillation when guided by direct energy surrogates rather than FLOPs can improve the energy consumption, sustainability, and deployability of LLMs for SE applications, enabling efficient models on consumer hardware.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑