发表机构
Southern University of Science and Technology(南方科技大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对LLMs在SPICE网表识别与操作中的可靠性,构建含2342个案例的NetlistBench基准,发现其性能随结构复杂度变化,为电路设计自动化提供关键瓶颈参考。
AI 中文摘要
大型语言模型(LLMs)正越来越多地被用于电路设计工作流中,然而它们在面向模拟器的SPICE网表识别与操作上的可靠性仍未被充分了解,且很少与高层设计推理分开考量。尽管网表是文本形式的,但它们通过拓扑结构和参数对结构化电路对象进行编码。我们提出NetlistBench,一个针对SPICE网表识别与操作的经结构验证的基准。NetlistBench包含2342个案例,覆盖24个任务族,涵盖参数与连接性识别和编辑、分层操作、等价性判断以及长时序复合编辑。模型输出由确定性的感知结构的预言机进行评估。在六个非思考型LLMs中,性能随操作级结构复杂性有显著变化:简单的局部编辑达到96%至100%的准确率,而器件添加降至41%至83%,等价性判断降至49%至90%。启用推理能显著提升较弱模型的性能,但无法消除结构保持失败,且随着编辑时序增加,性能仍会急剧下降。NetlistBench将网表可靠性确定为基于LLM的可信电路设计自动化的一个独特瓶颈。
英文摘要
Large Language Models (LLMs) are increasingly used in circuit design workflows, yet their reliability on simulator-facing SPICE netlist recognition and manipulation remains poorly understood and is rarely separated from high-level design reasoning. Although netlists are textual, they encode structured circuit objects through topology and parameters. We present \textbf{NetlistBench}, a structure-verified benchmark for SPICE netlist recognition and manipulation. NetlistBench contains 2,342 cases across 24 task families, covering parameter and connectivity recognition and edits, hierarchical operations, equivalence judgment, and long-horizon compound editing. Model outputs are evaluated by a deterministic structure-aware oracle. Across six non-thinking LLMs, performance varies substantially with operation-level structural complexity. Simple local edits reach $96\%$--$100\%$ accuracy, while device addition drops to $41\%$--$83\%$ and equivalence judgment to $49\%$--$90\%$. Enabling reasoning substantially improves weaker models but does not eliminate structure-preservation failures, with performance still degrading sharply as the edit horizon increases. NetlistBench identifies netlist reliability as a distinct bottleneck for trustworthy LLM-based circuit design automation.
Commentsaccepted by MLCAD 2026