发表机构
ETH Zurich(苏黎世联邦理工学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对CUTLASS内核选择,提出硬件感知特征表示,通过静态估计硬件行为增强配置,训练排序模型,显著降低选择遗憾,并验证跨精度迁移的有效性。
AI 中文摘要
GPU库(如CUTLASS)为单个操作暴露了数以万计语义等价的内核,使得穷举自动调优代价高昂,且无执行的选择变得困难。现有的分析型选择器需要手工设计的性能规则,而学习型选择器则基于原始配置参数操作,必须从数据中推断硬件后果。我们引入了一种面向CUTLASS内核选择的硬件感知表示,该表示用候选配置诱导的硬件行为的静态可计算估计来增强候选配置。我们构建了一个包含490万个CUTLASS内核的数据集,并训练梯度提升和神经学习排序模型,以对每个问题内的候选进行排序。在保留的穷举评估问题上,硬件感知表示相对于结构基线将选择遗憾降低了高达40%,相对于NVIDIA的矩阵乘法启发式降低了64.2%。我们进一步评估了CUTLASS GEMM内跨精度和尾声融合的数据高效迁移,表明显式表示候选诱导的硬件行为为学习型内核选择提供了有用的归纳偏置。
英文摘要
GPU libraries such as CUTLASS expose tens of thousands of semantically equivalent kernels for a single operation, making exhaustive autotuning expensive and execution-free selection difficult. Existing analytical selectors require hand-designed performance rules, while learned selectors operate on raw configuration parameters and must infer hardware consequences from data. We introduce a hardware-aware representation for CUTLASS kernel selection that augments candidate configurations with statically computable estimates of induced hardware behavior. We construct a dataset of 4.9 million CUTLASS kernels and train gradient-boosted and neural learning-to-rank models to rank candidates within each problem. On held-out exhaustive evaluation problems, hardware-aware representations reduce selection regret by up to 40\% relative to structural baselines and 64.2\% relative to NVIDIA's matrix-multiply heuristics. We further evaluate data-efficient cross-precision and epilogue-fusion transfer within CUTLASS GEMM, showing that explicitly representing candidate-induced hardware behavior provides a useful inductive bias for learned kernel selection.
Comments20 pages, 19 figures