ZhuLong:基于执行的大语言模型智能体,用于结合离线API自探索的EDA脚本编写
ZhuLong: Execution-Grounded LLM Agent for EDA Scripting with Offline API Self-Exploration
浏览论文内容
中文总结 AI 辅助
本文提出基于执行的LLM智能体ZhuLong,结合API检索、沙箱执行与离线API自探索机制,在EDA脚本任务基准测试中性能远超纯LLM基线,为EDA脚本编写提供了有效解决方案。
中文摘要 AI 辅助
针对工具特定且常无文档的EDA脚本API,是现有大语言模型(LLM)未能解决的长尾瓶颈。本文提出ZhuLong,一个用于PyAether和SKILL的基于执行的LLM编码智能体,它结合了API检索、文档检查和通过统一MCP工具实现的沙箱执行,并辅以离线API自探索机制,该机制通过反事实实验推断无文档API的行为。我们在EDA-Eval-PyAether上评估ZhuLong,这是一个包含158项真实任务的基准,采用基于断言的执行方式,完整系统在商业Empyrean Aether环境中实现了78.5%的Pass@1,显著优于纯LLM基线(23.6%)。消融研究表明,沙箱执行是性能的主要驱动因素(移除后性能下降41.2个百分点),而自探索机制额外贡献了3.2个百分点的准确率提升,并减少了22.1%的每任务工具调用次数。在涉及未保存布局和原理图的20项交互任务上,ZhuLong在PyAether上实现了60.0%的Pass@1,在SKILL上实现了50.0%的Pass@1。
英文摘要
EDA scripting with tool-specific, often undocumented APIs remains a long-tail bottleneck that existing LLMs fail to address. This paper presents ZhuLong, an execution-grounded LLM coding agent for PyAether and SKILL that combines API retrieval, documentation inspection, and sandbox execution via unified MCP tools, augmented by an offline API self-exploration mechanism that infers undocumented API behaviors through counterfactual experimentation. We evaluate ZhuLong on EDA-Eval-PyAether, a benchmark of 158 real-world tasks with assertion-based execution, where the complete system achieves 78.5% Pass@1 in the commercial Empyrean Aether environment, substantially outperforming a pure LLM baseline (23.6%). Ablation studies identify sandbox execution as the dominant performance driver (41.2 pp drop when removed), with the self-exploration mechanism contributing an additional 3.2 pp accuracy gain and a 22.1% reduction in per-task tool calls. On 20 interactive tasks involving unsaved layouts and schematics, ZhuLong achieves 60.0% Pass@1 for PyAether and 50.0% for SKILL.
发表机构
- Changxin Memory Technologies, Inc.(长鑫存储技术有限公司)
机构由 AI 辅助整理,请以论文原文为准。