发表机构
University of Minnesota Twin Cities; Simon Fraser University; Argonne National Laboratory(明尼苏达大学双城分校; 西蒙弗雷泽大学; 阿贡国家实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出HLSmith框架,结合HLS优化专业知识库、分阶段反馈驱动编排流程及模型适配流水线,在PolyBench上较ChatHLS实现4.24倍几何平均加速比,所有基准均生成正确设计,开放权重模型下最高加速比达138倍。
AI 中文摘要
专用FPGA加速器在众多应用领域能提供显著的性能和能效提升,但开发成本高昂,通常需要数月的专业工作。即便使用高级综合(HLS),设计人员仍需丰富的硬件专业知识才能构建高性能加速器。尽管大型语言模型(LLM)展现出强大的软件生成能力,但即便是前沿模型也缺乏将基础C/C++程序可靠转换为高性能HLS设计所需的硬件直觉和程序知识:它们难以识别有效架构、遵循HLS专家使用的优化流程,以及在不同内核间一致应用硬件转换。我们提出HLSmith,一个用于将C/C++程序转换为优化HLS加速器的专家引导框架。HLSmith包含三个组件:编码受保护转换规则、其适用性与前提条件及需避免的不安全情况的HLS优化专业知识库;以专家HLS开发实践为模型的分阶段、反馈驱动的编排流程,指导智能体完成综合、瓶颈分析和优化;以及基于工具的模型适配流水线,将商业前沿模型的优化轨迹转换为训练数据,用于微调开放权重LLM。我们在PolyBench上对HLSmith与ChatHLS(领先的现有HLS加速器开发智能体编排框架)进行评估。HLSmith在所有基准测试中均生成功能正确的设计(软件和RTL仿真均验证),几何平均加速比达ChatHLS的4.24倍,而ChatHLS的有效设计率仅为57%。此外,HLSmith使用商业前沿模型和开放权重模型时,分别实现最高252倍和138倍的加速比。
英文摘要
Application-specific FPGA accelerators offer substantial performance and energy-efficiency gains across many application domains, but developing them is costly, often requiring months of specialized effort. Even with high-level synthesis (HLS), designers still need extensive hardware expertise to build high-performance accelerators. Although large language models (LLMs) have demonstrated strong software-generation capabilities, even frontier models lack the hardware intuition and procedural knowledge needed to reliably translate baseline C/C++ programs into high-performance HLS designs: they struggle to identify effective architectures, follow the optimization processes used by HLS experts, and apply hardware transformations consistently across diverse kernels. We present HLSmith, an expert-guided framework for translating C/C++ programs into optimized HLS accelerators. HLSmith combines three components: an HLS optimization expertise library that encodes guarded transformation recipes, their applicability and prerequisite conditions, and unsafe cases to avoid; a staged, feedback-driven orchestration flow modeled on expert HLS development practice that guides agents through synthesis, bottleneck analysis, and optimization; and a tool-grounded model-adaptation pipeline that converts optimization trajectories from commercial frontier models into training data for fine-tuning open-weight LLMs. We evaluate HLSmith on PolyBench against ChatHLS, a leading prior agent-orchestration framework for HLS accelerator development. HLSmith achieves a geometric mean speedup of 4.24x over ChatHLS while producing functionally correct designs, in both software and RTL simulation, for every benchmark, compared with ChatHLS's 57% valid-design rate. It further reaches speedups of up to 252x and 138x with commercial frontier models and open-weight models, respectively.