AI 中文总结
研究针对GPU编程框架中性能分析工具滞后问题,提出TileSight工具,以瓦片为中心从核心到集群进行性能建模,能预测单GPU内核延迟、缓存命中率,在多GPU场景也有良好表现,优化时可选择有竞争力的瓦片配置。
AI 中文摘要
近期的GPU编程框架如Triton、TileLang和CUDA Tile将瓦片作为一等原语,使以瓦片为中心的编程成为高性能GPU内核的主流方法。但性能分析工具却未跟上,程序员仍依赖粗略的屋顶线边界、不透明的机器学习预测器或事后分析器来理解内核执行。对于现代人工智能工作负载,这种差距很明显,内核融合和分布式推理依赖张量核心、CUDA核心、缓存层次结构、内存管道和GPU间网络。我们提出TileSight,一个以瓦片为中心的性能建模工具,将瓦片从编程原语提升为分析原语。在GPU核心内,它对计算-内存管道重叠建模;跨核心时,对缓存层次结构建模;跨GPU时,对节点间通信建模。所有层共享瓦片抽象:瓦片内层级将工作表示为跨越网络、内存和计算管道的资源向量;瓦片间层级调度相关和有序操作以揭示合法重叠,并从瓦片重用距离推断多级缓存命中率;跨设备层级将远程张量访问映射到布局并通过α-β阶段成本进行路由。在A100、H200、B200和B6000上,TileSight预测单GPU内核延迟的合并平均绝对百分比误差(MAPE)为12.35%,优于现有基线且在不同架构间迁移性更好。其L2缓存命中率预测与各GPU上的测量值相差约一个百分点。在多达32个GPU时,TileSight在融合分布式内核上实现16.18%的加权MAPE(wMAPE),在端到端vLLM服务上实现13.52%的wMAPE。在优化方面,TileSight选择的瓦片配置可与强大的供应商和专家基线竞争。TileSight将在发表后开源。
英文摘要
Recent GPU programming frameworks such as Triton, TileLang, and CUDA Tile adopt tiles as first-class primitives, making tile-centric programming the prevailing approach for high-performance GPU kernels. Performance-analysis tooling has not followed: programmers still rely on coarse roofline bounds, opaque ML predictors, or post-hoc profilers to understand kernel execution. This gap is acute for modern AI workloads, where kernel fusion and distributed inference depend on tensor cores, CUDA cores, cache hierarchies, memory pipelines, and inter-GPU networks. We present TileSight, a tile-centric performance-modeling tool that elevates the tile from a programming primitive to an analysis primitive. Within a GPU core, TileSight models compute-memory pipeline overlap; across cores, it models the cache hierarchy; across GPUs, it models inter-node communication. All layers share the tile abstraction: the intra-tile layer expresses work as a resource vector spanning network, memory, and compute pipelines; the inter-tile layer schedules dependent and ordered actions to expose legal overlap and infers multi-level cache hit rates from tile reuse distance; and the cross-device layer maps remote tensor accesses to placements and routes them through an alpha-beta stage cost. On A100, H200, B200, and B6000, TileSight predicts single-GPU kernel latency with 12.35% pooled mean absolute percentage error (MAPE), outperforming state-of-the-art baselines and transferring better across architectures. Its L2 cache-hit-rate predictions are within roughly one percentage point of measurements on every GPU. At up to 32 GPUs, TileSight achieves 16.18% weighted MAPE (wMAPE) on fused distributed kernels and 13.52% wMAPE on end-to-end vLLM serving. In optimization, TileSight selects tile configurations competitive with strong vendor and expert baselines. TileSight will be open-sourced upon publication.