arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SyclKittens:面向Intel GPU的程序员与编码代理的瓦片编程模型

SyclKittens: A Tile Programming Model for Programmers and Coding Agents on Intel GPUs

Yehong Jiang, Sheng Chen, Fangwen Fu, Yen-Kuang Chen, Xinmin Tian, Stuart H. Sul, Simran Arora

arXiv 2610.04277首次发表:更新:

发表机构

Massachusetts Institute of Technology; Intel Corporation; Stanford University; California Institute of Technology(麻省理工学院; 英特尔公司; 斯坦福大学; 加州理工学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

SyclKittens是一种面向Intel GPU的硬件感知瓦片编程模型,通过编码已知高效方法,使编码代理能将核函数性能从57.5%提升至82.1%,并构建达到oneDNN约96%性能的核函数套件。

AI 中文摘要

新型AI加速器往往先于使其高效的核函数问世,因为峰值核函数性能需要操作数流水线和数据移动技术方面的架构特定专业知识。编码代理现在能够自行编写、编译和调优核函数,因此它们有望大幅加速核函数的开发与优化。代理所产出的结果取决于它们所获得的接口。这些接口可能暴露机器的原始能力,或编码利用其硬件特性高效运作的已知优良方法。我们通过三个编码模型在三个核函数上的受控实验,测试了这两种接口在Intel GPU上的性能影响。使用原始SYCL和执行反馈,三个测试模型中最强的Opus 4.8编写出的单GPU GEMM核函数仅达到Intel调优的oneDNN库的57.5%。我们提出SyclKittens,一种面向Intel GPU的硬件感知瓦片编程模型,它编码了这些已知优良方法,使其操作将操作数置于矩阵引擎布局中,通过L1缓存预取,并通过将GPU栈连接成节点的结构进行通信。在相同代理、任务和反馈预算下,使用SyclKittens的代理达到了82.1%,表明编码方法将硬件能力转化为性能。工作负载特定策略保持可编程性,工程师和代理在SyclKittens中共同细化这些调度,构建了一个在GEMM形状的几何平均值上达到oneDNN约96%的核函数套件。该套件在单GPU上运行Llama-3.1-8B推理比此http URL快1.59倍,在Intel的oneCCL上比匹配的多GPU解码路径快最多2.91倍。SyclKittens是开源的,可在以下https URL获取。

英文摘要

New AI accelerators arrive before the kernels that make them fast, because peak kernel performance requires architecture-specific expertise in operand pipelines and data-movement techniques. Coding agents can now write, compile, and tune kernels on their own, so they could greatly accelerate kernel development and optimization. What agents produce depends on the interfaces they are given. These interfaces may expose a machine's raw capabilities or encode the known-good methods for exploiting its hardware features efficiently. We test the performance impact of these two interfaces on Intel GPUs through controlled experiments with three coding models on three kernels. With raw SYCL and execution feedback, Opus 4.8, the strongest of the three models tested, writes a single-GPU GEMM kernel reaching only 57.5% of Intel's tuned oneDNN library. We present SyclKittens, a hardware-aware tile programming model for Intel GPUs that encodes these known-good methods, so its operations place operands in matrix-engine layouts, prefetch through the L1 cache, and communicate over the fabric that links GPU stacks into a node. With SyclKittens under the same agent, task, and feedback budget, the agent reaches 82.1%, showing that encoded methods turn hardware capabilities into performance. Workload-specific policies stay programmable, and engineers and agents jointly refine these schedules in SyclKittens to build a kernel suite that reaches ~96% of oneDNN in geometric mean across GEMM shapes. The suite runs Llama-3.1-8B inference 1.59x faster than torch$.$compile on one GPU and up to 2.91x faster than a matched multi-GPU decode path on Intel's oneCCL. SyclKittens is open source and available at https://github.com/intel/SyclKittens.

Comments46 pages, 11 figures. Code at https://github.com/intel/SyclKittens

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑