arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AsmEvo:采用功能等价性验证的AMD GPU内核智能体级汇编层优化工具

AsmEvo: Agentic Assembly-Level Optimization of AMD GPU Kernels with Functional Equivalence Verification

Ji Liu, Puyuan Yang, Rongzhang Zheng, Fan Wang, Jinglin Wang, Muhammad A. Awad, Mortis Huang, Andy Chang, Zekai Li, Zeping Li, Zihao An, Yue Liu, Yuchen Yang, Jianghui Wang, Chushi Chen, Ziqiong Liu, Fuwei Yang, Dong Li, Wen Heng Chung, Shengcai Liu, Emad Barsoum

arXiv 2608.20711首次发表:更新:

发表机构

AMD(超威半导体公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

AsmEvo是一款针对AMD GPU内核的智能体级汇编层优化工具,通过功能等价性验证优化已编译的AMDGPU代码对象,在多款AMD GPU上实现了显著加速比。

AI 中文摘要

高性能机器学习系统越来越依赖GPU内核,这些内核的可编辑源代码不可用、已生成或与最终机器代码距离过远,无法暴露剩余优化空间。现有的大语言模型(LLM)内核优化器和自动调谐器主要针对CUDA、Triton、HIP或张量程序源代码进行操作,并通过参考实现进行验证。本研究探讨更严格的场景:优化已编译的AMDGPU代码对象,其中部署的二进制文件是唯一的行为参照。我们提出AsmEvo,一款针对AMD GPU内核的智能体级汇编层优化工具。给定AMDGPU代码对象K0,AsmEvo会重建可重新汇编的表示形式,利用长视野智能体提出低级编辑,重建符合应用二进制接口(ABI)的优化对象,并仅在相同启动条件下与K0进行差分验证后才接受候选方案。AsmEvo结合了代码对象恢复、元数据感知重建、性能分析引导的热点窗口编辑、正确性门控计时以及保守的原位补丁回退机制。我们针对各类AMD GPU内核对AsmEvo开展了大量实验:在MI308X上,AsmEvo优化了30个选定KernelBench内核中的29个,实现了1.35倍几何平均加速比和3.88倍最大加速比;在MI300X生产工作负载上,其优化了所有评估的AITer二进制文件以及vLLM/SGLang的Triton汇编内核,分别实现了1.09倍/1.31倍和1.18倍/1.34倍的几何平均/最大加速比,同时保持功能等价性。

英文摘要

High-performance ML systems increasingly rely on GPU kernels whose editable source is unavailable, generated, or too distant from final machine code to expose remaining optimizations. Existing LLM kernel optimizers and autotuners mainly operate on CUDA, Triton, HIP, or tensor-program source and validate against reference implementations. We study a stricter setting: optimizing an already compiled AMDGPU code object, where the deployed binary is the only behavioral oracle. We present AsmEvo, an agentic assembly-level optimizer for AMD GPU kernels. Given an AMDGPU code object K0, AsmEvo reconstructs a reassemblable representation, proposes low-level edits with a long-horizon agent, rebuilds an ABI-preserving optimized object, and accepts candidates only after differential verification against K0 under identical launches. AsmEvo combines code-object recovery, metadata-aware rebuilding, profiling-guided hot-window editing, correctness-gated timing, and conservative in-place patch fallback. We conduct extensive experiments with AsmEvo on various AMD GPU kernels. On MI308X, AsmEvo improves 29 of 30 selected KernelBench kernels, reaching 1.35x geometric-mean and 3.88x maximum speedup. On MI300X production workloads, it improves all evaluated AITer binaries and vLLM/SGLang Triton assembly kernels, reaching 1.09x/1.31x and 1.18x/1.34x geometric-mean/maximum speedups, respectively, while preserving functional equivalence.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑