ForgeMegakernel:一种用于高效自回归模型解码Megakernel的通用框架
ForgeMegakernel: A General Framework for Efficient Auto-Regressive Model Decode Megakernels
- Tsinghua University(清华大学)
- University of Chinese Academy of Sciences(中国科学院大学)
- ModelBest Inc.(魔搭社区)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
提出ForgeMegakernel框架,利用编码代理和里程碑知识库及测试预言机自动生成高性能解码megakernel,在多个模型上实现显著加速。
AI中文摘要:
自回归模型解码受带宽限制,因为每个权重和键/值缓存字节每个令牌都会穿过高带宽内存一次。Megakernel是一种理想的解决方案,但现有的自动megakernel生成方法无法同时实现跨模型的泛化和正确性保证。我们提出了ForgeMegakernel,它使用编码代理为每个模型生成高性能解码megakernel。ForgeMegakernel将包含十个渐进式里程碑的通用知识库与独立的中间状态测试预言机配对。这些里程碑提供了megakernel的结构特性:每个SM的细粒度指令流、取代全局同步的依赖计数器,以及用于跨SM工作负载平衡和更高并行度的共享内存缓冲区池。测试预言机推导megakernel的中间状态,并在生成过程中检查性能、错误和精度,确保生成正确且可信的megakernel。我们在跨越0.6B-13B参数的八个模型家族的14种代表性解码操作上评估了ForgeMegakernel。生成的megakernel在相同配置下实现了50.5-85.9%的MBU,几何平均加速比相对于SGLang 0.5.18为1.21倍,相对于megakernel编译器为1.54倍。在SGLang内部,使用不规则提示在GSM8K上评估时,所有14个megakernel在可比答案准确度下解码速度均快于SGLang引擎。
英文摘要:
Auto-regressive model decode is bandwidth-bound, since every weight and key/value-cache byte crosses high-bandwidth memory once per token. A megakernel is an ideal solution, but existing automatic megakernel generation approaches cannot achieve both generalization across models and correctness guarantees. We present ForgeMegakernel, which generates a per-model high-performance decode megakernel using coding agents. ForgeMegakernel pairs a universal knowledge base of ten progressive milestones with an independent mid-state test oracle. The milestones provide the megakernel's structural properties: a fine-grained instruction stream for each SM, dependency counters replacing the global synchronization, and a shared-memory buffer pool for workload balance across SMs and greater parallelism. The test oracle derives the mid-states of the megakernel and checks the performance, error and precision during the generation process, guaranteeing a correct and trustworthy forged megakernel. We evaluated ForgeMegakernel on 14 representative decoding operations across eight model families spanning 0.6B-13B parameters. The generated megakernels achieved 50.5-85.9% MBU and geometric mean speedups of 1.21x over SGLang 0.5.18 and 1.54x over a megakernel compiler under identical configurations. Inside SGLang, evaluated on GSM8K with ragged prompts, all 14 megakernels decoded faster than the SGLang engine at comparable answer accuracy.