arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.03114cs.PL

用基于坐标的爬山法增强基于多面体的优化的能力

Enhancing the Power of Polyhedral-Based Optimizations with Coordinate-Based Hill Climbing

  • Stony Brook University(石溪大学)
  • Cadence(楷登电子)
  • UFMG(米纳斯吉拉斯联邦大学)

机构由 AI 辅助整理,请以论文原文为准。

Gaurav Verma, Michael Canesche, Fernando Magno Quintão Pereira

AI总结:

该研究将逐坐标爬山调优器扩展至多面体编译器Pluto,经调优的CPU内核性能优于默认配置及静态优化器,GPU线程块分配调优也有提升,搜索成本低于AutoTVM,为多面体编译与自动调优提供了折中方案。

AI中文摘要:

本文介绍了我们将多面体编译器Pluto扩展为轻量级逐坐标爬山调优器的实践,该调优器在Pluto选定内核的循环结构后,会调整诸如分块大小和线程块维度之类的数值变换参数。为确保快速收敛并跳出局部极小值,爬山法辅以两项技术:扩展邻域探索和最短跳细化阶段。在x86和ARM CPU上,经调优的内核性能优于Pluto的默认配置(11个基准测试的几何平均加速比为1.06-1.28倍),也优于静态优化器(Clang -O3、Polly、IOOpt),其性能可与AutoTVM自动调优器媲美,但搜索成本显著更低。将相同技术应用于NVIDIA A100上的GPU线程块分配,相比默认配置可获得5.5-8.5%的性能提升。这些结果表明,后优化参数调优是固定成本模型多面体编译与全自动调优之间的实用折中方案。

英文摘要:

This paper describes our experience extending the polyhedral compiler Pluto with a lightweight, coordinate-wise hill-climbing tuner that adjusts numeric transformation parameters, such as tile sizes and thread-block dimensions, after Pluto selects the kernel's loop structure. To ensure fast convergence and escape local minima, hill climbing is augmented with two techniques: expanded neighborhood exploration and a shortest-hop refinement phase. On x86 and ARM CPUs, tuned kernels outperform Pluto's default configuration (1.06-1.28x geometric mean speedup across 11 benchmarks) and static optimizers (Clang -O3, Polly, IOOpt), reaching performance competitive with the AutoTVM autotuner at substantially lower search cost. Applying the same technique to GPU thread-block allocation on an NVIDIA A100 yields 5.5-8.5% improvement over default configurations. These results position post-optimization parameter tuning as a practical middle ground between fixed-cost-model polyhedral compilation and full autotuning.

补充信息

↑