arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于大语言模型嵌入的程序分析与优化

LLM-Based Embeddings for Program Analysis and Optimization

Calvin Higgins, Marco Alvarez

arXiv 2608.07894首次发表:更新:

AI 中文总结

该研究首次将LLMCompiler生成的程序嵌入应用于程序分析与优化任务,结合源与IR代码嵌入在算法分类上错误率1.54%、较SOTA提升12%,为代码优化提供新方向。

AI 中文摘要

近期研究进展凸显了机器学习,特别是大语言模型(LLMs)在程序分析与优化领域的应用潜力。我们首次将LLMCompiler(一种在中间表示(IR)代码上进行大规模预训练的大语言模型)生成的程序嵌入,应用于代表性的程序分析与优化任务。我们采用一种简单方法直接从源代码和IR代码生成程序嵌入:将程序拆分为多个块,用预训练的LLMs对每个块独立嵌入,再将块嵌入聚合为单个程序嵌入。实验表明,结合源代码与IR代码的嵌入,在算法分类任务中错误率达1.54%,较当前最优水平提升12%,在异构设备映射任务中也取得了有竞争力的准确率。这些发现表明,训练具备性能感知能力的大语言模型以嵌入IR代码,或可在代码优化任务中达到最优结果。

英文摘要

Recent advances have highlighted the potential of machine learning, particularly Large Language Models (LLMs), for analyzing and optimizing programs. We present the first application of program embeddings from LLMCompiler---an LLM massively pretrained on intermediate representation (IR) code---to representative program analysis and optimization tasks. We generate program embeddings directly from source and IR code using a simple approach: split programs into chunks, independently embed each chunk with pretrained LLMs, and then aggregate the chunk embeddings into a single program embedding. Our experiments show that combining source and IR code embeddings achieves an error rate of 1.54\% in algorithm classification, a 12\% improvement over the current state-of-the-art, and a competitive accuracy on heterogeneous device mapping. These findings suggest that training a performance-aware LLM for embedding IR code might yield state-of-the-art results in code optimization tasks.

Journal ref2025 International Joint Conference on Neural Networks (IJCNN), Rome, Italy, 2025, pp. 1-8

DOI:10.1109/IJCNN64981.2025.11227194

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑