arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

MetaKernelBench:衡量超越代码的GPU内核知识迁移

MetaKernelBench: Measuring GPU Kernel Knowledge Transfer Beyond Code

Xueyi Chen, Shiyu Liu, Xin Jin, Yuhua Zheng, Xin Li, Haolei Bai, Junhan Zhu, Huan Wang

arXiv 2610.05014首次发表:更新:

发表机构

Westlake University; Shanghai Jiao Tong University; Zhejiang University(西湖大学; 上海交通大学; 浙江大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

MetaKernelBench通过跨DSL技能迁移测试,衡量GPU内核优化智能体的知识复用价值,发现迁移效果依赖于源尝试的相对表现。

AI 中文摘要

近期的GPU内核优化智能体将所学内容保留在知识库中或作为提炼后的技能。内核基准测试对每次尝试的实现进行正确性和速度评分,但未衡量保留经验的复用价值。我们引入MetaKernelBench,用于衡量从一种内核领域特定语言(DSL)中的一次尝试提炼出的经验,是否能在另一种DSL中改进对同一问题的新尝试。其74个问题为六个家族中的融合子图,每个问题以CuTe DSL和TIRx变体配对形式提出,两者仅在DSL上不同。智能体首先单独尝试每个变体,并被指示将所学内容提炼为自然语言技能,无论源尝试是否通过验证,该技能都会被迁移。该技能是同一模型在另一种DSL中进行技能条件化尝试的唯一额外输入。我们在匹配的每次尝试预算下,将每次技能条件化尝试与同一变体上的单独尝试进行比较,对正确性和端到端运行时间进行评分。在六个模型和两个方向上,相对于单独尝试的配对提升范围从-19%到+29%。四个模型在两个方向上均获得提升,但在每个模型和方向上,16%至45%的问题出现性能回退。结果遵循源尝试相对于目标单独尝试的结果,而非仅源尝试成功:当源尝试优于目标时,71%的比较中有所提升;当源尝试劣于目标时,54%的比较中发生回退。MetaKernelBench通过衡量同一问题的跨DSL内核知识迁移,补充了实现质量指标。

英文摘要

Recent GPU kernel optimization agents retain what they learn in knowledge bases or as distilled skills. Kernel benchmarks score each attempt's implementation for correctness and speed but leave the reuse value of retained experience unmeasured. We introduce MetaKernelBench, which measures whether experience distilled from an attempt in one kernel domain-specific language (DSL) improves a fresh attempt at the same problem in another. Its 74 problems are fused subgraphs in six families, each posed as a pair of CuTe DSL and TIRx variants that differ only in the DSL. The agent first attempts each variant solo and is instructed to distill what it learns into a natural-language skill, which is transferred whether or not the source attempt passes verification. The skill is the only extra input to a skill-conditioned attempt by the same model in the other DSL. We compare each skill-conditioned attempt with the solo attempt on the same variant under matched per-attempt budgets, scoring correctness and end-to-end runtime. Across six models and both directions, paired lift over solo attempts ranges from -19% to +29%. Four models gain in both directions, yet regressions occur on 16% to 45% of problems in every model and direction. Outcomes follow the source attempt's result relative to the target's solo attempt rather than source success alone, improving in 71% of comparisons when the source stands above and regressing in 54% when it stands below. MetaKernelBench complements implementation-quality metrics by measuring same-problem cross-DSL kernel knowledge transfer.

CommentsProject page: https://yige24.github.io/MetaKernelBench

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑