发表机构
George Mason University(乔治梅森大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出TileBench,一个在NVIDIA B200上评估Triton和cuTile的受控基准,包含45个算子,发现性能差距取决于工作负载,且Triton在LLM生成内核时更高效。
AI 中文摘要
基于瓦片的编程模型(如Triton和cuTile)旨在简化高性能内核开发,但其实际性能、调优行为和可用性仍难以系统比较。我们提出了TileBench,一个受控基准,用于在匹配的算子语义和可比较的实现结构下,在NVIDIA B200 GPU上评估Triton和cuTile。TileBench包含45个算子,涵盖多样化的AI内核模式和内存/计算行为。每个任务提供PyTorch参考、经过验证的Triton和cuTile实现、标准化的数据类型(dtype)和输入大小扫描、默认和自动调优配置、基于roofline的指标以及剖析引导的诊断。我们的评估表明,性能差距取决于工作负载:cuTile在少数Tensor-Core/TMA友好的内核上表现出色,而Triton在许多不规则、流式和带宽受限的算子上更强。我们进一步评估了LLM生成的cuTile和Triton内核,发现在相同的迭代细化协议下,Triton始终比cuTile更具token效率。TileBench在此https URL公开可用。
英文摘要
Tile-based programming models, such as Triton and cuTile, aim to simplify high-performance kernel development, but their practical performance, tuning behavior, and usability remain difficult to compare systematically. We present TileBench, a controlled benchmark for evaluating Triton and cuTile on NVIDIA B200 GPUs under matched operator semantics and comparable implementation structures. TileBench contains 45 operators covering diverse AI-kernel patterns and memory/computation behaviors. Each task provides a PyTorch reference, verified Triton and cuTile implementations, standardized data-types (dtype) and input-size sweeps, default and autotuned configurations, roofline-based metrics, and profiling-guided diagnosis. Our evaluation shows that performance gaps are workload-dependent: cuTile excels on a small cluster of Tensor-Core/TMA-friendly kernels, while Triton is stronger on many irregular, streaming, and bandwidth-bound operators. We further evaluate LLM-generated cuTile and Triton kernels and find that Triton is consistently more token-efficient than cuTile under the same iterative refinement protocol. TileBench is publicly available at https://github.com/Deep-Learning-Profiling-Tools/Tilebench.