arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SGAnalog:来自开源硅片流片的端到端电路基准

SGAnalog: An End-to-End Circuit Benchmark from Open-Source Silicon Tapeouts

Yueting Li, Weihang Ding

arXiv 2610.03934首次发表:更新:

发表机构

University of California, Berkeley(加州大学伯克利分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出基于Tiny Tapeout开源芯片的模拟电路基准,含273个拓扑设计,评估原理图到网表转录与器件尺寸调整,最强模型转录精确图同构率56.1%,尺寸调整达91.2分,揭示模型能力差异与策略限制。

AI 中文摘要

现有的模拟集成电路设计基准使得两个问题难以回答:一个模型是否学习了可迁移的电路技能,而非回忆熟悉的示例;以及其输出是否在定义的过程和测试条件下有效。我们引入了一个基准,该基准基于与Tiny Tapeout制造穿梭机相关的人工设计的开源电路构建。该集合包含273个拓扑上不同的顶层设计。每个源文件都在其穿梭机提交记录的修订版本中检索,并在固定的容器化环境中处理。流水线从同一源文件导出每个符合条件的原理图图像及其SPICE网表,为转录提供精确的结构参考。提交日期支持针对特定模型的训练截止分析,而作者提供的测试平台为尺寸调整提供了仿真环境。该基准评估原理图到网表的转录和设备尺寸调整。在固定66个转录任务的七个模型中,最强模型达到56.1%的精确图同构率,且七个模型中有六个从小规模到中等规模急剧下降。对于一个前沿模型,移除作者选择的标签会减少精确匹配,同时保持聚合结构F1,这表明标签有助于连接追踪。在17个尺寸调整任务中,领先模型在所有17个方案上收敛,并针对人类参考达到91.2分(满分100),而两个最新的Claude模型拒绝执行它们无异议转录的相同提示中的4个和11个;没有尺寸的方案得分为零。这两个任务产生不同的模型排名,揭示了不同的视觉和设计能力,以及在一个模型家族中,是策略而非能力限制。

英文摘要

Existing analog integrated circuit design benchmarks make two questions hard to answer: whether a model has learned transferable circuit skills rather than recalled familiar examples, and whether its output works under defined process and test conditions. We introduce a benchmark built from human-designed, open-source circuits associated with Tiny Tapeout manufacturing shuttles. The collection contains 273 topologically distinct top-level designs. Every source is retrieved at the revision recorded for its shuttle submission and processed in a fixed containerized environment. The pipeline exports each eligible schematic image and its SPICE netlist from the same source file, giving transcription an exact structural reference. Commit dates support model-specific training-cutoff analysis, while author testbenches provide the simulation context for sizing. The benchmark evaluates schematic-to-netlist transcription and device sizing. Across seven models on a fixed set of 66 transcription tasks, the strongest model reaches 56.1% exact graph isomorphism, and six of seven models drop sharply from the small to the medium tier. For one frontier model, removing author-chosen labels reduces exact matches while preserving aggregate structural F1, suggesting that labels can aid connectivity tracing. On the 17 sizing tasks, the leading model converges on all 17 proposals and reaches 91.2 out of 100 against the human reference, while the two newest Claude models refuse 4 and 11 of the same prompts they transcribe without objection; a proposal without sizes scores zero. The two tasks produce different model rankings, exposing distinct visual and design capabilities and, in one family, a policy rather than capability limit.

CommentsAccepted at the AI for Chip Design Workshop at NeurIPS 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑