从指标到改进:面向研究软件质量的生命周期感知大语言模型反馈框架
From Metrics to Improvement: A Lifecycle-Aware LLM Feedback Framework for Research Software Quality
AI总结:
针对研究软件质量问题,本文提出生命周期感知的LLM反馈框架,结合定量质量评估与迭代式LLM代码优化,实验验证其可改善代码重复等质量属性,同时揭示相关质量维度间的权衡关系。
AI中文摘要:
研究软件在科研工作流中的核心地位日益凸显,但通常由软件工程专业知识有限的研究人员开发,这会引发质量问题,阻碍软件的可维护性、可复现性、可复用性和可持续性。现有静态分析工具可识别此类问题,但其输出往往需要专业人员解读,且在将质量评估转化为可执行改进方面支持有限。为解决这一缺口,本文提出一种生命周期感知框架,将定量软件质量评估与基于大语言模型(LLM)的代码优化相结合。该框架包含两个阶段:第一阶段,基于成熟的软件质量标准和从业者需求,开发生命周期感知质量模型,该模型定义5个质量维度和25个候选指标,其中14个指标通过现有分析工具和自定义测量实现可操作化;第二阶段,将生成的质量诊断结果作为结构化反馈,用于迭代式的LLM优化过程,使生成的改进方案可根据质量模型反复重新评估。本文以以笔记本为核心的研究软件为对象,使用多个LLM对该框架进行评估,对比迭代式结构化反馈、单步反馈和非结构化提示的效果。结果显示,框架在特定质量属性(尤其是代码重复和结构质量)方面实现了改进,同时也揭示了可维护性、代码规模、文档和复杂性之间的权衡关系。这些发现证明了基于指标的LLM反馈在研究软件质量改进中的潜力,同时凸显了其固有的多目标特性。源代码和实验数据可在该公开链接获取。
英文摘要:
Research software is increasingly central to scientific workflows, yet it is often developed by researchers with limited software engineering expertise. This can lead to quality issues that hinder maintainability, reproducibility, reuse, and sustainability. Existing static analysis tools can identify such issues, but their outputs often require expert interpretation and provide limited support for translating quality assessments into actionable improvements. To address this gap, we propose a lifecycle-aware framework that integrates quantitative software quality assessment with Large Language Model (LLM)-based code refinement. The framework comprises two stages. First, a lifecycle-aware Quality Model is developed from established software quality standards and practitioner requirements. The model defines five quality dimensions and 25 candidate metrics, of which 14 are operationalized using existing analysis tools and custom measurements. Second, the resulting quality diagnostics are used as structured feedback within an iterative LLM-based refinement process, enabling generated improvements to be repeatedly reassessed against the Quality Model. We evaluate the framework on notebook-centric research software using multiple LLMs and compare iterative structured feedback with single-step feedback and unstructured prompting. The results show improvements in specific quality attributes, particularly code duplication and structural quality, while also revealing trade-offs among maintainability, code size, documentation, and complexity. These findings demonstrate the potential of metric-driven LLM feedback for research software quality improvement while highlighting its inherently multi-objective nature \footnote{The source code and experimental data are publicly available at https://github.com/QCDIS/Software_Quality_Control_LLM . }