arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.16831cs.DC

技术报告:NVIDIA Blackwell上的人工智能辅助门控DeltaNet优化

Technical Report: AI-Assisted Gated DeltaNet Optimization on NVIDIA Blackwell

Hyunjun Shin, Jiseung Jang, Jaewoo Maeng, Hyunjun Kim

首次发表
浏览论文内容

中文总结 AI 辅助

研究以MSInfer团队提交给MLSys 2026 FlashInfer竞赛的作品为例,探讨人工智能辅助GPU编程。该作品优化门控DeltaNet解码和预填充,在NVIDIA B200/Blackwell上实现1.58倍加速,揭示竞赛级优化需考虑多方面,将其视为端到端系统问题。

中文摘要 AI 辅助

人工智能辅助的GPU编程通常被构建为一个内核生成循环:要求模型生成更快的CUDA代码,对结果进行基准测试,然后重复。本案例研究认为,竞赛级别的优化不仅仅是改进内核主体。我们研究了我们团队MSInfer提交给MLSys 2026 FlashInfer竞赛的Agent-Assisted作品。该作品在NVIDIA B200/Blackwell上对门控DeltaNet解码和预填充进行了优化,实现了官方1.58倍的加速,解码的平均延迟约为9.315微秒,预填充的平均延迟约为239.4微秒。我们的经验表明,当工作负载需要结构重新制定和与评估器对齐的测量时,即使是有效的局部内核改进也可能达到瓶颈。因此,我们将人工智能辅助的内核优化描述为一个端到端的系统问题,它包括算法设计、工作负载专业化、测量工具、构建和评估表面、评估器对齐以及人工解释。

英文摘要

AI-assisted GPU programming is often framed as a kernel-generation loop: ask a model to produce faster CUDA code, benchmark the result, and repeat. This case study argues that contest-grade optimization involves more than improving the kernel body. We examine the Agent-Assisted submission by our team, MSInfer, to the MLSys 2026 FlashInfer Contest. The submission optimized Gated DeltaNet decode and prefill on NVIDIA B200/Blackwell and achieved an official $1.58\times$ speedup, with approximate average latencies of $9.315\,μ\mathrm{s}$ for decode and $239.48\,μ\mathrm{s}$ for prefill. Our experience shows that even effective local kernel improvements can plateau when a workload requires structural reformulation and evaluator-aligned measurement. We therefore characterize AI-assisted kernel optimization as an end-to-end systems problem that encompasses algorithm design, workload specialization, measurement tooling, build and evaluation surfaces, evaluator alignment, and human interpretation.

补充信息

↑