发表机构
Kyushu University; The University of Osaka(九州大学; 大阪大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出PreMaQ,利用LLM内部表示在代码生成前预测其可维护性质量(CSS和MI),在16个模型-基准组合上验证有效性,并用于模型选择以提升可维护性质量。
AI 中文摘要
随着大型语言模型(LLMs)在代码生成方面能力日益增强,在软件开发中采用生成的代码不仅需要评估其功能正确性,还需要评估其可维护性相关质量。如果这种质量能在生成之前被估计,开发者就可以避免生成、审查和丢弃低质量代码的成本。尽管先前的研究表明LLM生成代码的功能正确性可以提前预测,但可维护性相关质量是否同样可预测仍不清楚。我们引入了生成前可维护性相关质量预测(PreMaQ),该方法从LLMs的内部表示中在生成之前预测生成代码的代码坏味分数(CSS)和维护性指数(MI)。我们的评估涵盖了四个开放权重的LLMs和四个Python代码生成基准,共计2,695个任务。结果显示,预测的CSS和MI在所有16个模型-基准组合中与观测值持续相关,平均Spearman秩相关系数分别为0.57和0.65。当用于模型选择时,PreMaQ在多个模型生成功能正确代码的任务上,实现了理想的可维护性选择器相对于随机选择所能达到的可维护性相关质量改进的59.5%。将PreMaQ的预测与提示嵌入相结合,这一比例提高到60.7%,表明这两种信号是互补的。这些发现表明,PreMaQ可用于预测和改进LLM生成代码的可维护性相关质量。
英文摘要
As large language models (LLMs) become increasingly capable of code generation, adopting generated code in software development requires assessing not only its functional correctness but also its maintainability-related quality. If such quality could be estimated before generation, developers could avoid the cost of generating, reviewing, and discarding low-quality code. Although prior work has shown that the functional correctness of the LLM-generated code can be predicted in advance, it remains unclear whether maintainability-related quality is similarly predictable. We introduce Pre-Generation Maintainability-Related Quality Prediction (PreMaQ), which predicts the Code Smell Score (CSS) and Maintainability Index (MI) of generated code from the internal representations of LLMs before generation. Our evaluation covers four open-weight LLMs and four Python code generation benchmarks, comprising 2,695 tasks in total. Our results show that predicted CSS and MI consistently correlate with their observed values across all 16 model-benchmark combinations, achieving mean Spearman rank correlations of 0.57 and 0.65, respectively. When used for model selection, PreMaQ achieves 59.5% of the maintainability-related quality improvement attainable by an ideal maintainability-based selector over random selection on tasks for which multiple models generate functionally correct code. Combining predictions from PreMaQ and prompt embeddings increases this proportion to 60.7%, indicating that the two signals are complementary. These findings suggest that PreMaQ can be used to predict and improve the maintainability-related quality of LLM-generated code.
Comments10 pages, 4 figures