arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.26830cs.ARcs.AI

验证与仿真捕获不同错误:LLM生成电路的四个评估层级

Validation and Simulation Catch Different Errors: Four Levels of Evaluation for LLM-Generated Circuits

Ali Hedayati Pirouzan

首次发表
浏览论文内容

中文总结 AI 辅助

针对LLM生成电路,提出四个评估层级(模式、拓扑、可执行性、组件一致性),通过150电路基准和配对消融证明结构验证与仿真捕获不同错误,应独立报告。

中文摘要 AI 辅助

仿真成功并不等同于LLM生成电路的结构正确性。我们定义并测量了四个评估层级——模式有效性、拓扑有效性、后端可执行性和组件集一致性——在一个包含150个电路的三语基准上,通过一个基于类型化电路交换表示构建的部署流水线进行。这些层级并非嵌套关系。在gpt-4o-mini上,150个电路中有16个(10.7%,95%置信区间6.7-16.6)被拓扑验证器拒绝,但在ngspice中执行时无错误或警告;其中12个包含完全要求的组件,但有一个端子断开。相反,有7个电路(4.7%)通过了验证器但被ngspice拒绝。10个电路两项检查均失败,117个两项均通过,因此每项检查都能检测到另一项遗漏的类别。一个最小的三组件分压器展示了代价:一个悬空电阻报告5.00 V而非2.50 V,而ngspice保持沉默。一项配对消融实验,其中每个分支都从相同的模型样本而非新样本进行评估,将每个修复阶段与采样噪声分离。在一个分层的45电路子样本上,模型修复将拓扑有效性从40.0%提升至84.4%(+20个电路,无回退),同时将可执行性净提升6(+7,-1),该效应在此样本量下无法确定,组件一致性提升2。一个电路在单次修复步骤中在两个层级上朝相反方向变化。与直接网表基线相比,流水线执行了88.7%对比47.3%,或在一种将我们无法自信归因于网表的每次失败都计入基线的核算下为62.7%。这些结果支持一个狭窄的方法论结论:结构验证和仿真应作为LLM生成电路的独立评估阶段报告。能运行的电路不一定结构有效,结构有效的电路也不一定可执行。

英文摘要

Simulation success is not equivalent to structural correctness for LLM-generated circuits. We define and measure four evaluation levels -- schema validity, topological validity, backend executability, and component-set agreement -- on a 150-circuit trilingual benchmark, through a deployed pipeline built on a typed circuit interchange representation. The levels are not nested. On gpt-4o-mini, 16 of 150 circuits (10.7%, 95% CI 6.7-16.6) were rejected by the topological validator but executed in ngspice with no error or warning; 12 of these contained exactly the requested components, with one terminal disconnected. Conversely, 7 circuits (4.7%) passed the validator and ngspice refused them. Ten failed both checks and 117 passed both, so each check detects a class the other misses. A minimal three-component divider shows the cost: a dangling resistor reports 5.00 V instead of 2.50 V while ngspice stays silent. A paired ablation, in which every arm is evaluated from the same model sample rather than a fresh one, separates each repair stage from sampling noise. On a stratified 45-circuit subsample, model repair raised topological validity from 40.0% to 84.4% (+20 circuits, no regressions) while moving executability by a net 6 (+7, -1), an effect this sample size does not resolve, and component agreement by 2. One circuit moved in opposite directions at two levels in a single repair step. Against a direct-netlist baseline the pipeline executed 88.7% against 47.3%, or 62.7% under an accounting that credits the baseline with every failure we cannot confidently attribute to the netlist. These results support a narrow methodological conclusion: structural validation and simulation should be reported as distinct evaluation stages for LLM-generated circuits. A circuit that runs is not necessarily structurally valid, and a structurally valid circuit is not necessarily executable.

↑