arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.03983cs.PLcs.AI

大语言模型能否恢复编译器遗漏的语义优化机会?

Can Large Language Models Recover Semantic Optimization Opportunities That Compilers Miss?

Hailong Jiang, Feng Yu, Emran Hossain, Jianfeng Zhu, Mengfei Ren, Qiang Guan, Chunwei Xia

首次发表
浏览论文内容

中文总结 AI 辅助

本文探究大语言模型能否恢复编译器遗漏的语义优化机会,推出含多类案例的可执行基准SeGaBench,实验显示最强模型生成的工件可实现较高正确率与加速比,证实LLM可作为编译器的补充语义提议者。

中文摘要 AI 辅助

优化编译器会在被分析的程序表示中缺失可启用的语义时,遗漏可盈利的变换。本文探究大语言模型(LLM)能否从异构C/C++上下文恢复此类语义,并将其实现为经过验证、符合契约的工件。我们推出SeGaBench,这是一个可执行基准,包含100个合成案例和20个源自源代码的案例,涵盖低级假设、数据结构不变量以及高级语义提升。每个案例都包含隐藏的可启用语义、一个预言机工件、正确性和语义验证器,以及可复现的性能协议。我们对5个LLM进行评估,每个案例获取5个独立响应。表现最强的模型在94.8%的响应中生成正确工件,在83.3%的案例中实现至少1.05倍的加速,在93.3%的案例中获得性能成功。不过,正确的工件通常仅能填补预言机差距的一部分。这些结果表明,LLM可作为推测性语义提议者补充编译器分析,前提是其生成的工件经过验证和评估。

英文摘要

Optimizing compilers miss profitable transformations when their enabling semantics are absent from the analyzed program representation. We ask whether large language models (LLMs) can recover such semantics from heterogeneous C/C++ context and realize them as validated, contract-preserving artifacts. We introduce SeGaBench, an executable benchmark containing 100 synthetic and 20 source-backed cases spanning low-level assumptions, data-structure invariants, and high-level semantic lifting. Each case includes hidden enabling semantics, an oracle artifact, correctness and semantic validators, and a reproducible performance protocol. We evaluate five LLMs using five independent responses per case. The strongest model produces correct artifacts in 94.8% of responses, achieves at least 1.05x speedup in 83.3%, and obtains a performance success on 93.3% of cases. Nevertheless, correct artifacts often close only part of the oracle gap. These results show that LLMs can complement compiler analysis as speculative semantic proposers, provided that their artifacts are validated and evaluated.

发表机构

  • Youngstown State University(扬斯敦州立大学)
  • Kent State University(肯特州立大学)
  • Baylor University(贝勒大学)
  • University of Leeds(利兹大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑