发表机构
University at Buffalo; National Center for Supercomputing Applications; IBM Research(布法罗大学; 美国国家超级计算应用中心; IBM研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究评估了21种提示策略对LLM生成代码能效的影响,选出8种在10个LLM上测试,发现Python能耗最高降25%,C++最高降17%,凸显了提示设计对可持续编程的重要性。
AI 中文摘要
随着AI辅助编程日益主流,AI生成软件对环境的影响已成为重要考量。这促使在功能正确性之外,通过考虑执行效率和能耗来评估LLM生成的代码。然而,尽管代码生成取得了重大进展,前沿LLM很少根据其生成代码的能效进行评估。在本工作中,我们对21种用于能效代码生成的提示策略进行了全面评估,并确定了8种策略,在10个广泛使用的开源和专有LLM上进行评估。我们评估了这些策略在Python和C++代码生成中相对于基线提示的有效性。在评估的模型中,选定的提示策略在Python代码生成中实现了高达25%的能耗降低,在C++代码生成中实现了高达17%的能耗降低。在模型层面,Python能耗降低在Granite-4.0-H-Small上高达50%,在Claude 4.5 Haiku上高达39%,在MiniMax M3上高达28%,而C++能耗降低在Granite-4.0-H-Small上高达56%,在Qwen3-Coder-480B-A35B-Instruct上高达7%。这些结果表明,提示策略可以显著影响LLM生成代码的能耗,尽管其有效性因模型和编程语言而异。我们的发现强调了将能效纳入基于LLM的代码生成评估和优化的重要性,并为设计更可持续的AI辅助编程提示提供了实用见解。
英文摘要
As AI-assisted programming becomes increasingly mainstream, the environmental impact of AI-generated software has emerged as an important consideration. This motivates evaluating LLM-generated code beyond functional correctness by considering execution efficiency and energy consumption. However, despite substantial advances in code generation, frontier LLMs are rarely evaluated based on the energy efficiency of the code they produce. In this work, we conduct a comprehensive evaluation of 21 prompting strategies for energy-efficient code generation and identify 8 strategies for evaluation across 10 widely used open-weight and proprietary LLMs. We evaluate their effectiveness for both Python and C++ code generation relative to a baseline prompt. Across the evaluated models, the selected prompting strategies achieved energy reductions of up to 25% for Python and 17% for C++ code generation. At the model level, Python energy reductions reached up to 50% for Granite-4.0-H-Small, 39% for Claude 4.5 Haiku, and 28% for MiniMax M3, while C++ reductions reached up to 56% for Granite-4.0-H-Small and 7% for Qwen3-Coder-480B-A35B-Instruct. These results demonstrate that prompting strategies can substantially influence the energy consumption of LLM-generated code, although their effectiveness varies across models and programming languages. Our findings highlight the importance of incorporating energy efficiency into the evaluation and optimization of LLM-based code generation and provide practical insights into designing prompts for more sustainable AI-assisted programming.