arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SyntheticHLS:利用大语言模型构建多样化综合高级综合数据集

SyntheticHLS: Building Diverse Synthetic High-Level Synthesis Datasets using LLMs

Stefan Abi-Karam, Miaoyan Zhou, Callie Hao

arXiv 2610.00106首次发表:更新:

发表机构

Georgia Institute of Technology; Georgia Tech Research Institute(佐治亚理工学院; 佐治亚理工研究院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对HLS数据集稀缺且多样性不足的问题,提出SyntheticHLS框架,利用LLM的迭代反馈引导变异和定量指标生成多样化合成数据集,经交叉验证表明其训练语料库泛化性最佳。

AI 中文摘要

深度学习和大型语言模型(LLMs)正在半导体设计中迅速获得应用,推动了对训练数据集的需求。大多数工作集中在硬件描述语言(HDLs)上,而针对高级综合(HLS)——一种流行的领域专用加速器设计方法——的设计仍然稀缺。现有的HLS数据集工作强调人工整理或设计参数化,很少涉及基于LLM的高质量生成或代码长度、层次结构、设计空间大小、延迟、资源利用和应用领域的多样性,这可能限制模型的泛化能力。我们提出SyntheticHLS,一个利用LLMs生成大规模、复杂、多样化的合成HLS数据集的框架。其两个关键思想是:1)一个迭代的反馈引导变异循环,使用配对的HLS源代码和设计空间规范,逐步将种子设计转化为更复杂、可扩展的设计;2)HLS设计复杂性和设计空间可扩展性的定量指标,作为LLM引导变异的可测量目标。我们系统地交叉验证了一个HLS设计质量(QoR)深度学习模型,该模型在常见HLS基准、零样本合成设计和迭代变异合成设计上进行了训练和测试。合成设计在常见基准测试集上具有良好的迁移性,而反向迁移则不成立。SyntheticHLS的迭代变异设计在所研究的数据集中提供了最具泛化能力的训练语料库。对变异过程和数据集的分析表明,指标引导的轨迹持续改进目标复杂性和可扩展性目标,而不会使非目标指标退化。变异设计比零样本生成的设计覆盖了更广泛、更多样化的设计空间。我们的框架、数据集和评估是开源的:此 https URL。

英文摘要

Deep learning and large language models (LLMs) are rapidly gaining adoption in semiconductor design, driving demand for training datasets. Most efforts focus on hardware description languages (HDLs) while designs for high-level synthesis (HLS), a popular approach to domain-specific accelerators, remain scarce. HLS dataset efforts emphasize manual curation or design parameterization, seldom addressing high-quality LLM-based generation or diversity in code length, hierarchy, design-space size, latency, resource utilization, and application domain, potentially limiting model generalization. We propose SyntheticHLS, a framework for generating large-scale, complex, diverse synthetic HLS datasets using LLMs. Its two key ideas are: 1) an iterative feedback-guided mutation loop that uses paired HLS source code and design-space specifications to incrementally transform seed designs into more complex, scalable designs; and 2) quantitative metrics of HLS design complexity and design-space scalability that serve as measurable objectives for LLM-guided mutation. We systematically cross-validate an HLS Quality-of-Results (QoR) deep learning model trained and tested across common HLS benchmarks, zero-shot synthetic designs, and iteratively mutated synthetic designs. Synthetic designs transfer well to common benchmark test sets while the reverse does not hold. SyntheticHLS's iteratively mutated designs provide the most generalizable training corpus among the datasets studied. Analysis of the mutation process and dataset shows that metric-guided trajectories consistently improve targeted complexity and scalability objectives without regressing non-target metrics. Mutated designs span a substantially broader, more diverse design space than zero-shot generated designs. Our framework, dataset, and evaluation are open-source: https://github.com/sharc-lab/synthetic-hls.

CommentsAccepted and to be presented at the International Conference on Field Programmable Technology (FPT) 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑