发表机构
Princeton University; Meta(普林斯顿大学; Meta)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
Argus是一个以语义区域为中心的GPU性能测量规划器,自动编排跨层测量,在44个配置中改善39个,加速比从5.4%提升至8.9%,并显著提升内核优化与跨层PGO性能。
AI 中文摘要
GPU开发者和自动化优化器需要针对语义代码区域(如神经网络算子实现和流水线阶段)的性能证据,但这些证据分散在各种性能分析工具中。回答区域级问题可能需要手动构建探针和程序变体、隔离干扰测量,并将证据映射到区域和执行上下文。我们提出了Argus,一个以区域为中心的测量规划器和运行时,可自动化此工作流程。客户端通过边界标记识别区域,并选择信号和执行范围。Argus在编译、执行和测量变体中保持区域身份,构建干扰感知的多运行计划,并跨后端编排变换和性能分析。它利用区域身份和动态执行上下文连接编译器级、硬件级和系统级证据,生成记录测量来源和归因不确定性的报告。我们在智能体内核优化、持久化大内核优化和跨层PGO中评估了Argus。在44个持久化GEMM和注意力配置中,Argus改善了39/44个案例,并将AlphaEvolve的几何平均加速比从5.4%提升至8.9%。在持久化TinyLlama-1.1B解码大内核上,优化智能体使用Argus达到1.65毫秒/令牌,而未使用Argus时为4.92毫秒/令牌,生成的内核比使用CUDA Graphs的PyTorch快2.1倍。最后,Argus引导的跨层PGO改善了计算与通信的重叠,在五个多GPU设置中平均吞吐量提高了7%。
英文摘要
GPU developers and automated optimizers need performance evidence for semantic code regions--such as neural-network operator implementations and pipeline stages--but this evidence is fragmented across profiling tools. Answering a region-level question can require manually constructing probes and program variants, isolating interfering measurements, and mapping evidence to regions and execution contexts. We present Argus, a region-centric measurement planner and runtime that automates this workflow. Clients identify regions with boundary markers and select signals and execution scopes. Argus preserves region identity across compilation, execution, and measurement variants, constructs interference-aware multi-run plans, and orchestrates transformations and profiling across backends. It joins compiler-, hardware-, and system-level evidence using region identity and dynamic execution context, producing reports that record measurement origins and attribution ambiguity. We evaluate Argus across agentic kernel optimization, persistent megakernel optimization, and cross-level PGO. Across 44 persistent-GEMM and attention configurations, Argus improves 39/44 cases and raises AlphaEvolve's geometric-mean speedup from 5.4% to 8.9%. On a persistent TinyLlama-1.1B decode megakernel, an optimization agent reaches 1.65 ms/token with Argus versus 4.92 ms/token without it, producing a kernel $2.1\times$ faster than PyTorch with CUDA Graphs. Finally, Argus-guided cross-level PGO improves compute--communication overlap, increasing throughput by 7% on average across five multi-GPU settings.
Comments14 pages, 15 figures, 3 tables