arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.25831cs.SE

WarmTuner:通过离线到在线强化学习实现特定程序的编译器自动调优热启动

WarmTuner: Program-Specific Warm Starts for Compiler Autotuning via Offline-to-Online Reinforcement Learning

Tianlu Qiao, Mingxuan Zhu, Zeyu Sun, Dan Hao

首次发表
浏览论文内容

中文总结 AI 辅助

研究针对编译器自动调优中标志组合难寻的问题,提出WarmTuner框架,通过离线到在线强化学习,将历史记录转化为程序条件策略,结合GRPO在线更新,在评估中显著优于其他技术,平均加速1.732倍并在多个程序上获最佳结果。

中文摘要 AI 辅助

编译器是将高级程序转换为机器代码的基础软件工具。现代编译器有数百种优化,通过优化标志开启或关闭以提高生成代码的性能。但可能的标志组合数量呈指数增长,难以找到适合给定目标程序的标志配置。现有编译器自动调优技术通过修剪搜索空间、注入搜索偏差或预测配置性能来降低调优成本。然而,它们从历史数据中提取的知识在搜索开始后就固定了,运行时反馈仅指导搜索本身,而非先前的知识。当先前知识与目标程序不匹配时,这些方法会在搜索到良好配置之前浪费大量有限的在线预算。我们提出了WarmTuner,一种离线到在线的强化学习框架,它将历史记录转化为一个程序条件策略,该策略可预测整个标志空间中每个标志的设置,并在目标程序上保持适应性。离线时,WarmTuner从历史良好配置中学习整个标志空间上的程序条件策略。在线时,它使用实际编译运行反馈在目标程序上优化相同的策略,使策略由测量的加速驱动,而非限于历史数据。我们用Group Relative Policy Optimization(GRPO)实例化在线更新,该方法在同一轮中比较候选者并避免单独的值模型。我们在GCC 15.2.0上用cBench和PolyBench评估了WarmTuner。结果表明,WarmTuner比GCC -O3平均加速1.732倍,并在30个程序中的14个上获得最佳结果,显著优于比较技术。

英文摘要

Compilers are fundamental software tools that translate high-level programs into machine code. Modern compilers expose hundreds of optimizations, each turned on or off through an optimization flag, to improve the performance of the generated code. However, the number of possible flag combinations grows exponentially, making it difficult to find a flag configuration well suited to a given target program. Existing compiler auto-tuning techniques reduce tuning cost by pruning the search space, injecting search biases, or predicting configuration performance. Although some exploit program features, the knowledge they extract from historical data is frozen once search begins; runtime feedback then guides only the search itself, never the prior. As a result, when this prior mismatches the target program, these methods waste much of the limited online budget before the search reaches good configurations. We propose WarmTuner, an offline-to-online reinforcement learning framework that instead turns historical records into a program-conditioned policy that predicts each flag's setting over the full flag space and remains adaptable on the target program. Offline, WarmTuner learns this program-conditioned policy over the full flag space from historical good configurations. Online, it refines the same policy on the target program using real compile-run feedback, so that the policy is driven by measured speedups rather than limited to the historical data. We instantiate the online update with Group Relative Policy Optimization (GRPO), which compares candidates in the same round and avoids a separate value model. We evaluate WarmTuner on GCC 15.2.0 with cBench and PolyBench. The results show that WarmTuner achieves an average speedup of 1.732x over GCC -O3 and obtains the best result on 14/30 programs, significantly outperforming the compared techniques.

补充信息

↑