PyTorch 的硬件属性算子性能分析
Hardware-Attributed Operator Profiling for PyTorch
浏览论文内容
中文总结 AI 辅助
提出硬件归因流水线 Operator Profiler,自动关联硬件指标与 PyTorch 算子,在 RTX PRO 6000 上归因 95-100% 内核时间,并带来 1.76-2.24 倍加速。
中文摘要 AI 辅助
框架性能分析器在不使用硬件计数器的情况下暴露算子计时;GPU 性能分析器在不进行算子归因的情况下暴露硬件计数器。手动弥合这一差距容易出错且难以扩展。我们提出了 Operator Profiler,一个硬件归因流水线,通过三条互补的归因路径自动将硬件指标关联到 PyTorch 算子:CUPTI 相关性、带有每流区间树的 NVTX 时间包围,以及从调试工件中进行的 Inductor 融合映射增强。NVIDIA Nsight Compute (ncu) 硬件计数器通过调用顺序匹配与 NVIDIA Nsight Systems (nsys) 内核记录进行匹配,避免了跨不兼容时钟域的时间戳连接。一个精选的 20 计数器指标集,采用持续时间加权聚合,覆盖了所有硬件瓶颈轴;层去重将具有 N 层、K 个唯一结构类别的模型的 ncu 重放时间减少了 N/K 倍;GPU 时钟锁定控制用于算子级比较的内核持续时间聚合。在 NVIDIA RTX PRO 6000 Blackwell 上,Operator Profiler 对编译工作负载(GPT-2、SDPA Attention)的内核运行时的 95-100% 进行了归因;黑盒库后端(如 cuDNN RNN)被正确显示为超过 85% 未归因,而不是被静默丢弃。应用于基于性能分析的 FX 图优化时,归因性能分析在两个编译优化案例研究中产生了 1.76 倍至 2.24 倍的性能分析内核时间加速;第三个 LSTM 诊断案例将 cuDNN 重新分派识别为结构性修复,而非 FX 图重写。
英文摘要
Framework profilers expose operator timing without hardware counters; GPU profilers expose hardware counters without operator attribution. Bridging this gap manually is error-prone and does not scale. We present Operator Profiler, a hardware attribution pipeline that automatically links hardware metrics to PyTorch operators via three complementary attribution paths: PyTorch profiler CUPTI correlation, NVTX temporal enclosure with per-stream interval trees, and Inductor fusion- map enrichment from debug artifacts. NVIDIA Nsight Compute (ncu) hardware counters are matched to NVIDIA Nsight Systems (nsys) kernel records via invocation-order matching, avoiding timestamp joins across incompatible clock domains. A curated 20-counter metric set with duration-weighted aggregation covers all hardware bottleneck axes, layer deduplication reduces ncu replay time by a factor of N/K for models with N layers across K unique structural classes, and GPU clock locking controls the kernel-duration aggregates used for operator-level comparison. On an NVIDIA RTX PRO 6000 Blackwell, Operator Profiler attributes 95-100% of kernel runtime for compiled workloads (GPT-2, SDPA Attention); black-box library backends such as cuDNN RNN are correctly surfaced as greater than 85% unattributed rather than silently dropped. Applied to profile-guided FX graph optimization, attributed profiles yield 1.76x-2.24x profiled-kernel-time speedups on the two compiled optimization case studies; a third LSTM diagnostic case identifies cuDNN re-dispatch as a structural fix rather than an FX graph rewrite.
发表机构
- Yotta Labs(Yotta实验室)
机构由 AI 辅助整理,请以论文原文为准。