arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.12700cs.LGcs.ARcs.DC

面向大语言模型生成的GPU内核的契约级验证器,以及门控线性循环(GDN)族的原生Blackward反向传播算法

A Contract-Grade Verifier for LLM-Generated GPU Kernels, and a Native Blackwell Backward for the Gated-Linear-Recurrence Family

  • E3A Healthcare(E3A医疗公司)

机构由 AI 辅助整理,请以论文原文为准。

Rishi Shah, Rishav Shrestha

中文总结 AI 辅助

本文提出契约级GPU内核验证器,发现公开系统接受的2638个机器生成内核中39.5%存在无法修复的错误,还构建了GDN族的原生Blackwell反向传播算法,指出现有内核生成正确性信号被高估。

中文摘要 AI 辅助

利用语言模型生成GPU内核的系统报告了很高的正确率,但这些正确率来自单一的宽松测试:在固定形状的少量随机输入上运行内核,若输出与参考值接近则接受。内核可能通过该测试却仍存在隐性错误:它可能返回普通数值,而真实答案为NaN或无穷大;可能运行结果不一致;形状改变时会出错;或在fp16精度下累积误差,而参考值使用fp32精度。我们构建了能正确检查正确性的工具:包含12个对抗性门的契约级验证器,每个门都是正确内核必须满足的属性,其中多个门无容差,因此无法通过选择阈值来解释失败。面向外部,该验证器审计了2638个机器生成的内核,这些内核已被某公开系统的自身测试架接受为正确,结果发现39.5%的内核存在无法通过任何容差论证修复的错误,62.1%的内核至少存在一项违规。该领域的标准测试接受了1487个被验证器拒绝的内核,仅反过来拒绝了14个。我们通过四种独立方式验证该发现:7/7阳性对照、阈值校准扫描、与参考基准自身正确性代码的98.5%一致性,以及分层人工审计。面向内部,该验证器评估了我们自己的内核:门控线性循环(GDN)族的首个原生Blackwell tcgen05训练反向传播算法,包括该领域仍在回退模式下运行的反向状态阶段。我们通过双精度基准独立验证其正确性,并通过该算法训练了5个GDN族成员。内核生成报告进展背后的正确性信号远弱于数据显示的程度,而一组无容差的契约将弥合大部分差距。

英文摘要

Systems that generate GPU kernels with language models report high correctness rates. Those rates come from a single loose test: run the kernel on a few random inputs at one fixed shape and accept it if the output is close to a reference. A kernel can pass that test and still be silently wrong. It can return an ordinary number where the true answer is a NaN or an infinity, differ from run to run, break when the shape changes, or accumulate in fp16 where the reference keeps an fp32 total. We build the instrument that checks correctness properly: a contract-grade verifier of twelve adversarial gates, each a property a correct kernel must satisfy, several of them tolerance-free, so no choice of threshold can explain a failure away. Aimed outward, the verifier audits 2,638 machine-generated kernels that a public system's own harness had already accepted as correct. It finds 39.5% broken beyond any tolerance argument and 62.1% carrying at least one violation. The field's standard test accepts 1,487 kernels the verifier rejects, against only 14 the other way. We defend the finding four independent ways: a 7/7 positive control, a threshold-calibration sweep, 98.5% agreement with the reference benchmark's own correctness code, and a stratified hand-audit. Aimed inward, the verifier judges a kernel of our own: the first native Blackwell tcgen05 training backward for the gated-linear-recurrence (GDN) family, including the reverse-state stage the field still runs on a fallback. We establish its correctness independently, against a double-precision oracle, and train five family members through it. The correctness signal behind reported progress in kernel generation is far weaker than the numbers suggest, and a set of tolerance-free contracts would close most of the gap.

补充信息

↑