AI 中文总结
研究针对代码优化中现有策略引导方法的局限,提出MoST框架,整合多知识源,通过跨源跨场景聚类识别策略并转移,利用新算法和程序,在基准测试和实际项目中显著提升代码优化效果,优于其他方法。
AI 中文摘要
自动化代码优化通过重构源代码来提高程序性能,近期研究使用大语言模型生成优化补丁。最新方法是策略引导的,从历史优化提交中总结策略作为静态分析规则,用于匹配代码位置让大语言模型优化。但现有方法有局限:无法利用其他知识源的策略,且策略适用场景单一。为此提出MoST框架,它整合多知识源,统一表示证据对象,跨源跨场景聚类识别策略并按需转移到目标场景生成规则。采用新算法和程序,实验表明在基准测试和实际项目中,MoST比其他方法生成更多高质量补丁,性能提升显著。
英文摘要
Automated code optimization improves program performance by refactoring source code, and recent studies use LLMs to generate optimization patches. The newest approaches are strategy-guided: they summarize strategies from historical optimization commits as static analysis rules, and use these rules to match code locations for LLMs to optimize. However, these approaches have two limitations: (1) the strategies may come from other knowledge sources, such as textbooks and web pages, but the existing approaches cannot utilize them; (2) a strategy may be applicable to different scenarios, e.g., different programming languages, but existing approaches can only formalize strategies for the scenario to which the source commit belongs. To address these limitations, we propose MoST, an LLM-based code optimization framework that integrates multiple knowledge sources across scenarios. MoST uniformly represents items in different knowledge sources as evidence objects, clusters them in a cross-source and cross-scenario manner to identify strategies, and transfers them to the target scenario when necessary for generating static analysis rules. To implement this process, MoST employs a novel self-balanced weighted clustering algorithm to balance evidence objects from different knowledge sources, and a novel example transfer procedure to ensure the quality of the generated rules when transferring across scenarios. On a benchmark containing 151 C/C++, 150 Python, and 50 Rust historical optimization tasks, compared with SemOpt, MoST yields 24.44%-180.00% and 21.88%-37.50% more patches that are exactly the same as or semantically equivalent to developer patches, respectively. When optimizing 15 real-world projects, MoST achieves 19.72%-717.42% maximum improvements and 4.44%-258.17% average improvements for the performance tests in the projects, significantly outperforming SemOpt and Codex.