Harness Engineering for LLM-Driven GPU Kernel Generation
用于大语言模型驱动的GPU内核生成的工具工程
机构 * Baidu, Inc.(百度公司)
AI总结 在MLSys 2026 FlashInfer竞赛中,针对NVIDIA Blackwell B200 GPU,提出以工具为中心的大语言模型驱动的GPU内核优化系统,分离评估工具与优化控制器,利用Codex等生成候选内核,实验显示优化后平均延迟加速显著,且代理辅助内核效果更佳。
Comments 24 pages, 6 figures. Extended technical report on our submission to the MLSys 2026 FlashInfer AI Kernel Generation Contest. Code: https://github.com/syhya/mlsys26-flashinfer-contest