arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.06620cs.SE

理解软件重构的能源影响:一项针对受控示例与真实提交的工作负载感知研究

Understanding the Energy Impact of Software Refactoring: A Workload-Aware Study of Controlled Examples and Real-World Commits

Haibo Wang, Heng Li, Shin Hwei Tan

AI总结:

该研究通过微基准与真实世界基准的大规模实证分析,发现重构的能源影响受工作负载显著影响,现有方法无法可靠识别重构引发的能源回归,需开发更准确的预测技术。

AI中文摘要:

重构可在保留功能行为的同时提升软件可维护性,但行为保留并不意味着能源中性。现有研究主要在固定或简单工作负载下考察孤立的重构操作,对工作负载变化、真实世界重构实践、解释因素及能源回归识别的影响认识不足。我们开展了首个针对重构能源影响的大规模实证研究,涵盖两个互补的Java基准:包含68种重构类型并在多样工作负载下评估的微基准(Micro-benchmark),以及包含来自430个GitHub项目的481次真实世界重构提交的实用基准(Practical-benchmark)。通过重复配对能源测量,我们分析了工作负载敏感性、重构模式、解释因素,以及基于指标和大语言模型(LLM)的回归识别效果。在微基准中,384个重构-工作负载对里有199个(51.8%)呈现统计显著的能源差异,45.3%的重构实例在不同工作负载下改变了能源影响分类;在实用基准中,仅36次提交(7.5%)显示显著能源变化,尽管三分之二的提交差异至少达10%。仅重构类型不足以预测能源结果,而某些重复出现的重构组合与能源降低相关。执行时间变化始终能解释受控基准中的能源变异,但与真实提交中的能源变化相关性较弱。我们的研究结果强调需针对重构的能源影响开展多样工作负载评估;现有基于指标的方法和基于LLM的预测器均无法可靠识别重构引发的能源回归,推动了更准确的重构能源影响预测技术的发展。

英文摘要:

Refactoring improves software maintainability while preserving functional behavior, yet behavior preservation does not imply energy neutrality. Existing studies primarily examine isolated refactorings under fixed or simple workloads, leaving the effects of workload variation, real-world refactoring practices, explanatory factors, and energy regression identification insufficiently understood. We present the first large-scale empirical study of the energy impact of refactoring across two complementary Java benchmarks: a Micro-benchmark, comprising 68 refactoring types evaluated under diverse workloads, and a Practical-benchmark, containing 481 real-world refactoring commits from 430 GitHub projects. Using repeated paired energy measurements, we analyze workload sensitivity, refactoring patterns, explanatory factors, and the effectiveness of metric- and LLM-based regression identification. In the Micro-benchmark, 199 of 384 refactoring-workload pairs (51.8%) exhibit statistically significant energy differences, and 45.3% of refactoring instances change energy-impact classification across workloads. In the Practical-benchmark, only 36 commits (7.5%) show significant energy changes, although two-thirds differ by at least 10%. Refactoring type alone is insufficient to predict energy outcomes, while certain recurring refactoring combinations are associated with energy reductions. Changes in execution time consistently explain energy variation in the controlled benchmark but correlate weakly with energy changes in real-world commits. Our findings highlight the need for workload-diverse evaluation of the energy impact of refactoring; neither existing metric-based approaches nor LLM-based predictors can reliably identify refactoring-induced energy regressions, motivating the development of more accurate techniques for predicting the energy impact of refactoring.

补充信息

↑