发表机构
National University of Singapore; Meta(新加坡国立大学; Meta)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
Xtrace通过二进制级指令拼接实现高保真GPU内核内追踪,近零编译干扰,最小化运行时开销,显著优于现有工具,并提升LLM内核性能优化效率。
AI 中文摘要
现代GPU内核将越来越多的工作融合到单个内核中,内核内追踪已成为分析这些内核的主流方法。追踪会在内核中插入探针以记录其运行时状态,而追踪的保真度决定了性能优化的效率。不幸的是,现有工具在编译前插入探针。这些工具会干扰编译器的优化,因此它们追踪的二进制文件与GPU实际执行的二进制文件不同。它们还增加了显著的运行时开销。Xtrace是首个具有近零编译时干扰和最小化运行时开销的GPU内核追踪系统。Xtrace直接将探针插入编译后的内核二进制文件中。它仅重用插入地址处持有死值的寄存器,并利用编译器的危险表解决所有危险。它进一步调度指令顺序、寄存器分配和控制位,以最小化探针引入的运行时开销。Xtrace支持19种NVIDIA和AMD GPU架构,并在此https URL公开可用。我们在主要的生产级大语言模型(LLM)内核上,将Xtrace与最先进的追踪器NVIDIA的Neutrino和IKET进行了评估。在H100、B300和MI300X GPU上,Xtrace保留了内核指令的94-98%,而现有工具仅保留了8-48%。Xtrace仅增加0.9-2.8%的开销,而现有工具增加了3.8-75.6%。Xtrace引导编码代理以比现有追踪少3.9倍的迭代次数达到相同的FlashAttention-3性能。得益于我们的二进制级插桩,Xtrace还能追踪更快的闭源cuDNN内核,这引导代理将开源FlashAttention-4的吞吐量提升5.2-13.3%。
英文摘要
Modern GPU kernels fuse increasingly more work into a single kernel, and intra-kernel tracing has become the mainstream method to profile them. Tracing inserts probes into the kernel to record its runtime states, and the fidelity of the trace determines the efficiency of performance optimization. Unfortunately, existing tools insert probes before compilation. These tools interfere with the compiler's optimizations, so they trace a different binary from the one the GPU executes. They also add significant runtime overhead. Xtrace is the first GPU kernel tracing system with near-zero compile-time interference and minimized runtime overhead. Xtrace inserts probes directly into the compiled kernel binary. It reuses only the registers that hold dead values at the insertion address and resolves all hazards with the compiler's hazard tables. It further schedules the instruction order, register allocation, and control bits to minimize the runtime overhead the probe introduces. Xtrace supports 19 NVIDIA and AMD GPU architectures, and is publicly available for use at https://g-watch.github.io. We evaluate Xtrace on major production large language model (LLM) kernels against the state-of-the-art tracers Neutrino and IKET from NVIDIA. On H100, B300, and MI300X GPUs, Xtrace preserves 94-98% of the instructions of the kernel, while existing tools preserve only 8-48%. Xtrace adds only 0.9-2.8% overhead, while existing tools add 3.8-75.6%. Xtrace guides a coding agent to reach the same FlashAttention-3 performance with 3.9x fewer iterations than existing traces do. Thanks to our binary-level instrumentation, Xtrace also traces the faster closed-source cuDNN kernel, which guides the agent to lift the open-source FlashAttention-4 by 5.2-13.3% in throughput.
Comments13 pages of main text plus an appendix (24 pages in total), 16 figures, 5 tables. Xtrace is publicly available at https://g-watch.github.io