AI 中文总结
本文提出长时程模拟设计基准,含50个晶体管级任务,评估15种智能体配置,发现电气收敛是主要挑战,提供任务匹配拓扑可显著提升性能。
AI 中文摘要
编码智能体现已能维持数小时、工具驱动的循环,但它们在将长时程模拟与混合信号电路推进至电气规格达标方面的能力仍未得到测量。我们提出模拟设计基准(Analog Design Bench),这是一个由17位芯片设计者贡献的、包含50个晶体管级设计任务的长时程智能体基准。智能体使用开源仿真器工作,而一个隔离的验证器则基于规格的电气测试对所提交的电路进行评估。我们在2250次两小时的尝试中评估了15种智能体配置,观察到完全规格通过率从8.0%到78.0%不等。编码基准性能与模拟结果相关,但无法解释大部分性能差异。我们的失败分析表明,大多数不成功的提交没有记录合法性拒绝,但未通过电气验收,从而确定电气收敛是主要的端点挑战。我们测试了时间、推理努力、智能体框架和提供的设计知识作为干预措施。更长的预算和更高的推理努力能提升性能,而通用技能文档带来的益处甚微,有时甚至降低性能。提供与任务匹配的参考拓扑——一种理想化的电路知识产权检索形式——使DeepSeek V4 Pro提升了18.7个百分点,并主要加速了GPT-5.6 Sol。
英文摘要
Coding agents now sustain hours-long, tool-driven loops, yet their ability to carry long-horizon analog and mixed-signal circuits to electrical specification remains unmeasured. We introduce Analog Design Bench, a long-horizon agentic benchmark of 50 transistor-level design tasks contributed by 17 chip designers. Agents work with an open-source simulator, while an isolated verifier evaluates the submitted circuit using specification-based electrical tests. We evaluate 15 agent configurations across 2,250 two-hour attempts and observe full-specification pass rates from 8.0% to 78.0%. Coding-benchmark performance correlates with analog results but leaves much of the performance spread unexplained. Our failure analysis shows that most unsuccessful submissions have no recorded legality rejection but fail electrical acceptance, identifying electrical closure as the dominant endpoint challenge. We test time, reasoning effort, agent harness, and supplied design knowledge as interventions. Longer budgets and higher reasoning effort improve performance, while general skill documents provide little benefit and sometimes reduce performance. Supplying a task-matched reference topology, an idealized form of circuit-IP retrieval, raises DeepSeek V4 Pro by 18.7 percentage points and mainly accelerates GPT-5.6 Sol.
Comments25 pages, 10 figures, 6 tables