arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.12629cs.LG

CAKE:面向前沿核函数演进的编译器-智能体协同设计

CAKE: Compiler-Agent Co-Design for Frontier Kernel Evolution

Zihao Ye, Yingyi Huang, Hongyi Jin, Bohan Hou, Junru Shao, Zhongming Yu, Jinqi Chen, Meghan Cowan, Shiyi Cao, Shanli Xing, Hanfeng Chen, Vinod Grover, Tianqi Chen, Luis Ceze

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出CAKE编译器-智能体协同设计方案,通过硬件显式IR实现GPU核函数演进,在Flash-KMeans等基准测试中性能优于CUDA/PTX,相关成果已提交上游PR。

中文摘要 AI 辅助

GPU核函数智能体与GPU编程语言各自发展,导致专家级核函数难以复现。智能体通常将编译器视为固定黑盒,仅接收错误、正确性结果和时序信息;现有领域特定语言(DSL)要么隐藏关键调度决策,要么通过复杂布局抽象暴露这些决策。本文提出CAKE,一种编译器-智能体协同设计方案,其中智能体编写CAKE中间表示(IR),这是一种带类型、硬件显式的调度表示。CAKE暴露 warp(线程束)角色、内存移动、同步和流水线,同时支持验证、代价建模和本地化诊断。该框架自身不断演进:反复出现的故障会转化为验证器规则、IR原语、模型校准和可复用优化策略。在匹配实现的隐藏式Flash-KMeans在B200上的干净启动实验中,8000万token预算下的最优CAKE IR候选方案性能达到调优后FlashML基准的1.144倍,而直接CUDA/PTX方案仅为0.928倍。超出该基准,智能体生成的Kimi Delta Attention相较于官方FlashKDA实现,几何平均加速比达2.05倍,并通过端到端服务验证。由调度器支持的KNN和KMeans在400余种形状上的性能提升幅度为1.42倍至2.12倍,且已有4项核函数变更作为上游拉取请求(PR)提交。CAKE针对从安培(Ampere)到布莱克韦尔(Blackwell)的NVIDIA GPU,将单形状演进与库泛化及调度相分离。

英文摘要

GPU kernel agents and GPU programming languages have advanced separately, leaving expert kernels difficult to reproduce. Agents usually treat the compiler as a fixed black box and receive only errors, correctness outcomes, and timing, while existing DSLs either hide critical scheduling decisions or expose them through difficult layout abstractions. We present CAKE, a compiler-agent co-design in which agents author CAKE IR, a typed, hardware-explicit schedule representation. CAKE exposes warp roles, memory movement, synchronization, and pipelines while supporting verification, cost modeling, and localized diagnostics. The harness itself evolves: recurring failures become verifier rules, IR primitives, model calibrations, and reusable optimization tactics. In matched implementation-hidden Flash-KMeans clean starts on B200, the best CAKE IR candidate at an 80-million-token budget runs at 1.144x the tuned FlashML baseline, compared with 0.928x for direct CUDA/PTX. Beyond this benchmark, agent-generated Kimi Delta Attention achieves a 2.05x geometric-mean speedup over official FlashKDA and passes end-to-end serving validation. Dispatcher-backed KNN and KMeans improve performance by 1.42x to 2.12x across more than 400 shapes, and four kernel changes are available as upstream PRs. CAKE targets NVIDIA GPUs from Ampere through Blackwell and separates single-shape evolution from library generalization and dispatch.

发表机构

  • Carnegie Mellon University(卡内基梅隆大学)
  • NVIDIA(英伟达公司)

机构由 AI 辅助整理,请以论文原文为准。

↑