发表机构
Georgia Institute of Technology; Georgia Tech Research Institute(佐治亚理工学院; 佐治亚理工研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文扩展HLS-Eval基准,引入基于mini-swe-agent的智能体评估流程,初步结果显示开源LLM能解决所有简单HLS任务,需提高基准难度并分析轨迹以推动智能体化HLS设计。
AI 中文摘要
大型语言模型(LLMs)和AI智能体正越来越多地被探索用于硬件设计,包括高层数字设计。虽然大多数工作针对硬件描述语言(HDLs)的代码生成和编辑,我们先前的工作引入了HLS-Eval,一个用于评估LLMs在高层综合(HLS)设计任务上表现的开源基准。然而,那些评估侧重于零样本生成和编辑,未涉及智能体如何完成HLS设计任务。因此,我们基于开源mini-swe-agent框架扩展了HLS-Eval,增加了智能体化评估流程。该流程允许HLS设计智能体使用文件编辑工具,调用C++编译器进行自我验证,并在推理过程中迭代优化设计,同时记录智能体轨迹以分析成本、令牌使用量和迭代次数。我们在现有HLS-Eval基准上展示了初步结果。在初步评估中,我们发现开源LLMs配合智能体化框架能解决我们评估中所有简单的HLS代码生成任务,这强调了随着模型能力提升,需要扩大基准难度。通过分析成功和失败运行的轨迹,我们展示了模型规模、令牌使用量和轨迹长度与设计通过率之间的关系。这些结果为智能体化HLS设计奠定了基础,并随着模型能力的进步,推动了更难基准和新智能体工具的开发。
英文摘要
Large language models (LLMs) and AI agents are increasingly explored for hardware design, including high-level digital design. While most work targets code generation and editing for hardware description languages (HDLs), our prior work introduced HLS-Eval, an open-source benchmark for evaluating LLMs on high-level synthesis (HLS) design tasks. Those evaluations, however, focused on zero-shot generation and editing, leaving open how agents achieve HLS design tasks. We therefore extend HLS-Eval with an agentic evaluation flow built on the open-source mini-swe-agent framework. The flow lets HLS design agents use file-editing tools, invoke a C++ compiler for self-verification, and iteratively refine designs during inference, while logging agent traces for analysis of cost, token usage, and iteration count. We present initial results on the existing HLS-Eval benchmarks. In our initial evaluation, we find open-source LLMs paired with an agentic harness solve every simple HLS code generation task in our evaluation, underscoring the need to expand benchmark difficulty as model capabilities advance. Analyzing traces from passing and failing runs, we show how model size, token usage, and trajectory length relate to design pass rates. These results establish a foundation for agentic HLS design and motivate harder benchmarks and new agentic tooling as model capabilities progress.
CommentsPresented at the Architecture 2.0 workshop at ISCA 2026