arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.03971cs.ARcs.AI

基于LLM的HLS修复的可执行基准:设计复杂度与修复欠约束

An Executable Benchmark for LLM-Based HLS Repair:Design Complexity and Repair Underconstraint

Maisha Mastora, Dean Sullivan

首次发表
浏览论文内容

中文总结 AI 辅助

本研究构建首个可执行的HLS修复基准,评估四个LLM在125个实例上的修复能力,发现设计复杂度主导难度,并引入解多重性度量修复欠约束,前沿模型显著提升修复率。

中文摘要 AI 辅助

使用大型语言模型(LLM)自动修复高层次综合(HLS)设计是一个新兴但尚未充分探索的问题。虽然基于LLM的修复在寄存器传输级(RTL)Verilog上显示出强劲结果,但先前唯一一项针对HLS逻辑修复的系统性研究报告称GPT-4的修正准确率仅为10.5%,且未分析修复失败的原因或影响难度的因素。我们首次对基于LLM的HLS修复进行了全面评估,涵盖四个模型(GPT-4o、GPT-4o-mini、GPT-5.4和Claude Opus 4.6),在125个基准实例上跨越八个逻辑错误类型,这些实例来自三个开源HLS套件(CHStone、MachSuite、Polybench)。我们构建了首个可执行的APR风格HLS修复基准,配备套件特定的功能预言机,支持pass@k评估,而非先前工作中使用的字符串匹配近似。设计上下文和规模(而非仅错误类型)主导修复难度:修复率在复杂密码学内核(CHStone)上为6-45%,在紧凑算法内核(MachSuite)上为81-93%,尽管错误类型分布相同。我们引入解多重性(solution multiplicity),即重复修复尝试中生成的不同补丁的比例,作为修复欠约束的经验度量,并表明它在所有四个模型中强烈预测修复失败(Spearman rho = -0.393,p = 6.61 x 10^-20)。前沿模型将欠约束实例从41-46个(GPT-4o、GPT-4o-mini)减少到仅3个(Claude Opus 4.6),pass@1从42-53%提高到73-76%。SHFT错误在所有模型中始终难以修复,而语义错误类别(如缓冲区索引)仅在前沿规模下变得可靠可修复。这些发现表明,未来的APR基准必须包含复杂、可执行且弱可识别的设计,这些设计在前沿LLM修复后仍具挑战性。

英文摘要

Automated repair of High-Level Synthesis (HLS) designs using large language models (LLMs) is an emerging but underexplored problem. While LLM-based repair shows strong results on register-transfer level (RTL) Verilog, the only prior systematic study of HLS logic repair reports just 10.5% correction accuracy for GPT-4, with no analysis of why repair fails or what drives difficulty. We present the first comprehensive evaluation of LLM-based HLS repair across four models (GPT-4o, GPT-4o-mini, GPT-5.4, and Claude Opus 4.6) on 125 benchmark instances spanning eight logic bug types across three open-source HLS suites (CHStone, MachSuite, Polybench). We construct the first executable APR-style HLS repair benchmark with suite-specific functional oracles, enabling pass@k evaluation rather than the string-match approximations used in prior work. Design context and scale, rather than bug type alone, dominate repair difficulty: repair rates range from 6-45% on complex cryptographic kernels (CHStone) to 81-93% on compact algorithmic kernels (MachSuite) despite identical bug type distributions. We introduce solution multiplicity, the fraction of distinct patches generated across repeated repair attempts, as an empirical measure of repair underconstraint, and show it strongly predicts repair failure across all four models (Spearman rho = -0.393, p = 6.61 x 10^-20). Frontier models reduce underconstrained instances from 41-46 (GPT-4o, GPT-4o-mini) to just 3 (Claude Opus 4.6), with pass@1 improving from 42-53% to 73-76%. SHFT bugs remain consistently hard across all models, and semantic bug classes such as buffer indexing become reliably repairable only at frontier scale. These findings show that future APR benchmarks must include complex, executable, and weakly identifiable designs that remain challenging after frontier LLM repair.

发表机构

  • University of New Hampshire(新罕布什尔大学)

机构由 AI 辅助整理,请以论文原文为准。

↑