发表机构
National University of Singapore; Nanyang Technological University; CAIR, VinUniversity(新加坡国立大学; 南洋理工大学; VinUniversity 人工智能研究中心)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对 LLM 生成 GPU 内核的基准测试预言机,提出变异分析作为充分性度量,注入万余故障评估检测率,揭示官方检查器漏检 16.9% 故障,并优化测试套件提升至 98% 检测率。
AI 中文摘要
针对 LLM 生成的 GPU 内核的基准测试,仅通过少量随机输入和宽松的浮点容差来决定正确性,其判定结果如今已用于排行榜和强化学习奖励。近期研究一致认为这些检查器较弱,并通过人工修补——增加输入分布、模糊测试配方、收紧容差——但无法衡量任何修补是否足够。我们引入变异分析作为内核基准测试预言机的充分性度量:确定性规则向 188 个 KernelBench 问题的已验证 CUDA 实现中注入 10,303 个可编译故障,其中 7,384 个具有独立的杀死见证;任何测试协议按其检测到的比例评分。官方检查器确定性地漏掉了六分之一(16.9%)的有见证故障,且漏检按类别存在偏差:算术故障漏检 8.7%,但精度故障漏检 78.6%。该度量解释了原因(随归约规模增长的容差盲带;由合法浮点方差设定的输入激进性测量上限),审计了现有最强修补(KernelBench-Verified 的改进分解为隐藏输入贡献 +4.0 分和更紧容差贡献 +4.5 分,这一分解其作者无法计算),并揭露了一个已发表的模糊测试配方错误拒绝了正确内核 107 次。基于杀死矩阵优化测试套件,每个问题使用两个输入即可达到 98.0% 的检测率(保留集上 94.8%),且该度量的故障分类法对测试生成器的教导胜过原始故障本身。在 48 个完整架构中,盲区随规模增长而扩大,集中于深层同质流水线,且有两个问题被证明无法仲裁:其官方参考违反了基准自身针对 fp64 的容差。我们将所有内容作为 KernelBench-M 发布。
英文摘要
Benchmarks for LLM-generated GPU kernels decide correctness with a few random inputs and a loose floating-point tolerance, and their verdicts now feed leaderboards and reinforcement-learning rewards. Recent work agrees these checkers are weak and patches them by hand---extra input distributions, fuzzing recipes, tighter tolerances---with no way to \emph{measure} whether any patch suffices. We introduce mutation analysis as an adequacy metric for kernel-benchmark oracles: deterministic rules inject 10{,}303 compilable faults into verified CUDA implementations of 188 KernelBench problems, 7{,}384 of them with an independent kill witness; any test protocol is scored by the fraction it detects. The official check misses \textbf{one in six} witnessed faults (16.9%), deterministically, and the misses are skewed by family: 8.7% of arithmetic faults escape, but 78.6% of precision faults do. The metric explains why (a tolerance blind band growing with reduction size; a measured ceiling on input aggressiveness set by legitimate floating-point variance), audits the strongest existing patch (KernelBench-Verified's gain splits into $+4.0$ points from hidden inputs and $+4.5$ from tighter tolerance, a split its authors could not compute), and exposes a published fuzzing recipe that rejects \emph{correct} kernels 107 times. Optimizing suites over the kill matrix reaches 98.0% detection with two inputs per problem (94.8% held-out), and the measurement's fault taxonomy teaches a test generator more than the raw faults themselves. Across 48 whole architectures, the blindness grows with scale, concentrating in deep homogeneous pipelines, and two problems prove unrefereeable: their official references violate the benchmark's own tolerance against fp64. We release everything as \href{https://huggingface.co/datasets/Elfsong/KernelBench-M}{KernelBench-M}.