arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

通过MCP工具调用对硬件设计自动化的AI智能体进行基准测试

Benchmarking AI Agents for Hardware Design Automation via MCP Tool Calling

Leonardo Liparulo, Francesco Pierri

arXiv 2608.26199首次发表:更新:

发表机构

Politecnico di Milano(米兰理工大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究构建含MCP服务器的硬件设计自动化基准,评估7种开源AI智能体,发现其可靠性依赖任务结构与配置,为本地LLM智能体部署提供实用指导。

AI 中文摘要

我们探究由本地部署的大语言模型驱动的AI智能体,能否在符合行业实际的工具调用场景中可靠地自动化执行专家定义的硬件设计工作流。在这类场景中,工程师需通过专用工具执行重复性、依赖有序的操作,例如创建组件、添加端口和连接布线。组件规格的保密性约束及命名约定常禁止使用托管的专有API,因此需采用本地部署的模型。为研究该场景,我们构建了一个Model Context Protocol(MCP,模型上下文协议)服务器,其可复现嵌入式系统开发中使用的专有硬件设计工具的状态与依赖逻辑,并构建了涵盖单操作编辑、多步骤依赖链、无效请求、拼写错误提示及多服务器工具上下文的基准测试集。我们评估了7种开源模型,对比了系统提示、工具描述细节、上下文范围、单智能体与多智能体架构等流水线选择。结果显示,强大的模型可在基准测试工作流上实现接近完整的预期调用覆盖率,但可靠性高度依赖任务结构与智能体配置;全面的工具描述可持续减少失败,少样本提示会导致部分模型出现严重弃权(不执行),累积上下文会损害受约束的模型,多智能体分解虽能助力弱工作者或长会话,但需以额外调用为代价。这些发现为在有状态的硬件设计环境中部署本地LLM智能体提供了实用指导。

英文摘要

We ask whether AI agents powered by locally deployed large language models can reliably automate expert-defined hardware design workflows in an industry-realistic tool-calling setting. In these environments, engineers issue repetitive, dependency-ordered operations---such as creating components, adding ports, and wiring connections---through specialised tools. Confidentiality constraints on component specifications and naming conventions often preclude hosted proprietary APIs, motivating the use of locally deployed models. To study this setting, we build a Model Context Protocol (MCP) server that reproduces the state and dependency logic of a proprietary hardware design tool used in embedded system development and construct a benchmark covering single-operation edits, multi-step dependency chains, invalid requests, misspelled prompts, and multi-server tool contexts. We evaluate seven open-source models comparing pipeline choices including system prompts, tool-description detail, context scope, and single-agent versus multi-agent architectures. Results show that strong models can achieve near-complete expected-call coverage on the benchmarked workflows, but reliability depends strongly on both task structure and agent configuration. Comprehensive tool descriptions consistently reduce failures, few-shot prompting can cause severe inaction for some models, cumulative context harms constrained models, and multi-agent decomposition helps weak workers or long sessions at the cost of additional calls. These findings provide practical guidance for deploying local LLM agents in stateful hardware design environments.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑