发表机构
University of California, Irvine; HPE Labs(加州大学尔湾分校; 惠普企业实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出结构增强LLM PRISM,通过注入AST、CFG和DFG表示提升HLS编译指示优化,零样本下合成内核数达Llama3-8B的3.5倍,设计速度提升2.31倍。
AI 中文摘要
编译指示(pragma)的插入决定了高层次综合(HLS)设计的质量。选择正确的指令需要专家知识以及对循环嵌套、数据依赖和内存布局的推理。虽然现有的大型语言模型(LLM)在代码生成方面显示出潜力,但它们缺乏显式的程序结构感知,限制了其提出有效编译指示的能力。我们提出了PRISM,一种新颖的结构增强型LLM,通过向预训练的、冻结的代码LLM添加编译器级别的结构推理来弥补这一差距。它结合了三种层次化的程序表示——抽象语法树(AST)、控制流图(CFG)和数据流图(DFG),并将它们注入到特定的Transformer层,同时将原始代码令牌保留在单独的流中。注入点的交叉注意力门允许在结构信号无帮助时回退到预训练表示。在HLS-Eval的零样本评估中,PRISM综合出的内核数量是Llama3-8B的3.5倍(26.9%对比7.7%),并且在成功的内核上,它生成的设计比GPT-5-mini快2.31倍(几何平均)。在智能体流程中,PRISM代码生成在优化复杂代码时优于其他基线,并将HLS-Eval套件的平均归一化改进提升至26.4%。
英文摘要
Pragma insertion drives the quality of high-level synthesis (HLS) designs. Choosing the right directives demands expert knowledge and reasoning about loop nesting, data dependences, and memory layout. While existing large language models (LLMs) show promise in code generation, they lack explicit program-structure awareness, limiting their ability to suggest effective pragmas. We present PRISM, a novel structure-augmented LLM that closes this gap by adding compiler-grade structural reasoning to a pretrained, frozen code LLM. It combines three hierarchical program representations, Abstract Syntax Tree (AST), Control-Flow Graph (CFG), and Data-Flow Graph (DFG), injecting them into a specific transformer layer while keeping original code tokens in a separate stream. The cross-attention gate at the injection point allows falling back to the pretrained representation when its structural signal is unhelpful. On zero-shot evaluation in HLS-Eval, PRISM synthesizes 3.5\times as many kernels as Llama3-8B (26.9\% vs. 7.7\%), and on the kernels where it does succeed, it produces designs that are 2.31\times faster (geomean) than GPT-5-mini's. In the agentic flow, the PRISM codegen outperforms other baselines when optimizing complex code and drives the average normalized improvement across the HLS-Eval suite to 26.4\%.
Journal ref2026 ACM/IEEE International Symposium on Machine Learning for CAD (MLCAD '26), September 07--09, 2026, Jeju Island, Republic of Korea