arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

张量内核的算子感知混合精度容差校准

Operator-Aware Mixed-Precision Tolerance Calibration for Tensor Kernels

Dipankar Sarkar

arXiv 2607.16228首次发表:更新:

发表机构

Arizona State University(亚利桑那州立大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究张量内核正确性测试的容差校准问题,通过挖掘测试用例误差分布提出新容差,经实验,对特定错误变体校准容差提高了错误检测召回率,虽有少量误报增加,但整体提升了检测效果。

AI 中文摘要

大多数张量内核正确性测试通过固定形状的全接近式检查,使用手动挑选的绝对和相对容差。这些阈值在整个语料库中复制且很少重新审视。我们从26个条目组成的gpuemu语料库和2种数据类型(8076个结果行)的累积云GPU运行中挖掘每个测试用例的逐元素误差分布。然后提出一个实证问题:在正确实现下内核本身能接受的绝对容差是多少?答案比当前手动挑选的绝对容差要严格得多。最大收紧是attention_triton fp16达到2184倍。对于语料库中提供配对正确对应物的七个LLM风格错误变体,按(算子,数据类型)校准容差将错误检测召回率从73.2%(2467个中的1805个)提高到82.4%(2467个中的2034个),绝对增益9.3个百分点(新增229个检测)。控制误报计数从1882个正确控制案例中的0个增加到20个(增加1.1个百分点)。

英文摘要

Most tensor-kernel correctness tests go through a fixed-shape all close-style check with hand-picked absolute and relative tolerances. The thresholds are copied across the corpus and rarely revisited. We mine the element-wise error distribution of every test case from accumulated cloud GPU runs across the 26-entry gpuemu corpus and 2 dtypes (8,076 result rows). We then ask one empirical question: what absolute tolerance would the kernel itself, observed under its correct implementation, justify? The answer is much tighter than the current hand-picked atol. The largest tightening is attention_triton fp16 at $2{,}184\times$. Restricted to the seven LLM-style buggy variants for which the corpus ships a paired correct counterpart, calibrated per-(op, dtype) tolerances raise bug-detection recall from 73.2% (1,805 of 2,467) to 82.4% (2,034 of 2,467), an absolute gain of 9.3 percentage points (+229 new detections). The control false-positive count rises from 0 to 20 out of 1,882 correct-control cases (+1.1 percentage points).

Comments8 pages, 1 figure, LNCS format. Companion paper: arXiv:2606.20128 (P1)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑