加速基于迁移学习的自动调优:利用预测性LLVM IR性能排序
Accelerating Transfer-Learning-Based Autotuning with Predictive LLVM IR Performance Ranking
- Iowa State University(爱荷华州立大学)
- Clemson University(克莱姆森大学)
- AMD(超威半导体公司)
- Argonne National Laboratory(阿贡国家实验室)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
提出基于迁移学习的自动调优框架,利用神经配置评分器对LLVM IR性能排序,减少评估开销,实现与最先进技术相当的性能,评估次数最多减少61.67%。
AI中文摘要:
随着高性能计算(HPC)生态系统的复杂性不断增加,实现最优性能成为一项挑战。传统的性能自动调优技术为应对这种复杂性提供了有前景的手段,但这些技术计算密集,且需要大量评估才能找到最优配置。本工作提出了一种自动调优框架,设计了一个基于机器学习的集成LLVM中间表示(IR)排序器——神经配置评分器(NCS)。NCS对由基于迁移学习的自动调优器采样的IR性能进行排序,通过减少调优开销并规避低质量评估来提高调优过程的效率。通过利用相关任务的知识,我们能够有效利用迁移关系,以比依赖迭代细化的传统技术更少的样本访问高性能配置。我们的框架可以实现与最先进的自动调优技术相似的性能提升,同时评估次数最多减少61.67%,在各种HPC基准测试中平均减少27.85%的评估次数。
英文摘要:
As the complexity of High Performance Computing (HPC) ecosys- tems continually increases, achieving optimal performance becomes a challenge. Traditional performance autotuning techniques pro- vide promising means to navigate this complexity, these techniques remain computationally intensive and require many evaluations to find optimal configurations. This work proposes an autotuning framework that designs a machine learning-based ensemble LLVM Intermediate Representa- tion (IR) ranker, Neural Configuration Scorer (NCS). NCS ranks the performance of IRs sampled by a transfer-learning-based autotuner, improving the efficiency of the tuning process by reducing tuning overheads and circumventing subpar evaluations. By leveraging knowledge from related tasks, we are able to effectively exploit the transfer relationship to access high-performing configurations in fewer samples than traditional techniques that rely upon itera- tive refinement. Our framework can achieve similar performance improvements as state-of-the-art autotuning techniques with up to 61.67% fewer evaluations, averaging 27.85% fewer evaluations across various HPC benchmarks.