arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

不要重新生成,要调试:一种用于修复近失配硬件算子的领域特定智能体

Don't Regenerate, Debug: A Domain-Specific Agent for Repairing Near-Miss Hardware Operators

Yansong Sun, Shenxiu Wu, Siyuan Chen, Runlin Hou, Junhao Qiu, Junming Cao, Shudi Shao, Zhichao Lu, Qingfu Zhang

arXiv 2608.02712首次发表:更新:

发表机构

City University of Hong Kong; Huawei Technologies Ltd.(香港城市大学; 华为技术有限公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究提出领域特定调试智能体,通过调试而非重新生成修复硬件近失配算子,其调试Pass@1准确率更高、token消耗更少,可拓展能力边界并降低成本。

AI 中文摘要

针对GPU、NPU等硬件加速器的内核生成已成为大型语言模型(LLM)的试验场,最先进的系统通过将LLM与智能体强化学习、进化搜索相结合的流水线提升正确性。这类流水线会生成、编译并执行大量候选内核,丢弃其中大部分,放弃了将失败提炼为可复用知识的机会。许多被丢弃的候选是近失配算子,它们可编译并运行,但数值验证失败;每个都蕴含真实领域知识,且在LLM推理、交叉编译、硬件执行上投入了非平凡成本。我们主张范式转变:不要重新生成,要调试。调试比从头生成受约束得多:搜索空间小,反馈密集。我们提出一种领域特定调试智能体,解决自主修复的三个核心挑战:通过检索模式和诊断工具缓解知识稀缺;通过反作弊检测和全覆盖评估确保完整性;通过收敛保护和有界迭代控制成本。调试有两个互补作用:它通过恢复反复重新生成无法产生的算子拓展能力边界,且降低每个可交付算子的成本。调试Pass@1达到66.7%,而重新生成平均Pass@1为25.9%、重新生成Pass@3为40.7%;每次成功的token消耗比三次尝试重新生成少92.8%。组件消融实验显示,知识库驱动恢复,完整性门拒绝了工作流本身接受的12.5%-33.3%的成功案例。

英文摘要

Kernel generation for hardware accelerators such as GPUs and NPUs has become a proving ground for large language models (LLMs), and state-of-the-art systems raise correctness through pipelines that couple LLMs with agentic reinforcement learning and evolutionary search. Such pipelines generate, compile, and execute large numbers of candidate kernels, discarding most of them and forgoing the opportunity to distill failures into reusable knowledge. Many discarded candidates are near-miss operators that compile and run but fail numerical validation; each embodies genuine domain knowledge and a nontrivial investment in LLM inference, cross-compilation, and hardware execution. We argue for a paradigm shift: rather than regenerate, debug. Debugging is far more constrained than generating from scratch: the search space is small and feedback is dense. We present a domain-specific debug agent that addresses three core challenges in autonomous repair: mitigating knowledge scarcity through retrieved patterns and diagnostic instrumentation, ensuring integrity through anti-cheat detection and full-coverage evaluation, and controlling cost via convergence guards and bounded iteration. Debugging serves two complementary roles: it extends the capability frontier by recovering operators that repeated regeneration fails to produce, and it lowers cost per deliverable operator. Debug Pass@1 achieves 66.7% versus Regenerate Avg Pass@1's 25.9% and Regenerate Pass@3's 40.7%, while consuming 92.8% fewer tokens per success than three-trial regeneration. Component ablations show that the knowledge base drives recovery, while integrity gates reject 12.5-33.3% of the successes the workflow itself accepted.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑