arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

2D-FET-Bench:从空间推理到薄片上的场效应晶体管设计

2d-fet-bench: from spatial reasoning to fet design on flakes

Dunzhi Zhou, Chengyu Zhu, Gang Qiu, Caiwen Ding

arXiv 2610.07423首次发表:更新:

发表机构

University of Minnesota; Future Intelligence Labs(明尼苏达大学; 未来智能实验室)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对二维薄片FET布局手工绘制问题,提出2D-FET-Bench V2基准,含128个任务,验证语言模型智能体构造能力,最佳配置GPT5.6-Luna ReAct-3实现80.5%覆盖率,证明任务可解并评估几何结构布局。

AI 中文摘要

在剥离的二维薄片上,场效应晶体管(FET)布局通常需要针对每个薄片手工绘制,根据其在光学显微图像中的位置和轮廓放置接触电极和栅极。据我们所知,目前尚无可执行的基准测试来检验语言模型智能体能否可靠地完成这种特定于薄片的构造任务。我们推出了2D-FET-Bench V2,这是一个包含128个布局任务的基准测试,这些任务基于显微镜衍生的薄片轮廓构建,包括含孔薄片和多薄片任务。每个任务都提供文本形式的器件规格说明和轮廓坐标。智能体生成类型化的多边形和路径操作,并渲染为GDSII格式。一个确定性验证器检查几何和结构要求,另一个完整性检查则验证所提供的轮廓保持不变。脚本化的参考布局通过了全部128个任务,表明每个任务都是可解的。我们评估了六个模型以及GPT5.6-Luna的七种工作流和脚手架变体,每个任务尝试五次。在六模型面板中表现最佳的配置是采用ReAct-3的GPT5.6-Luna,其尝试通过率为62.3%,至少一次解决任务的比例(覆盖率)为80.5%,五次尝试全部通过的比例(一致性)为43.8%。ReAct-3在pass@1指标上比单次Plan-and-Execute高出27.0个百分点,而令牌消耗是其2.46倍。专家对每个已覆盖任务中一个抽样验证通过的布局进行了审计,涉及五种ReAct-3配置,接受率介于56.4%至63.5%之间。该基准测试评估了几何和结构性的FET布局构造能力。

英文摘要

Field-effect transistor (FET) layouts on exfoliated two-dimensional flakes are typically drawn by hand for each flake, placing contacts and gates to match its position and outline in optical micrographs. To our knowledge, no executable benchmark tests whether language-model agents can perform this flake-specific construction reliably. We introduce 2D-FET-Bench V2, a benchmark of 128 layout tasks built from microscopy-derived flake contours, including hole-containing flakes and multi-flake tasks. Each task supplies a textual device specification and contour coordinates. An agent generates typed polygon and path operations rendered to GDSII. A deterministic verifier checks geometric and structural requirements, and a separate integrity check verifies that the supplied contours remain unchanged. Scripted reference layouts pass all 128 tasks, showing that every task is solvable. We evaluate six models and seven workflow and scaffold variants of GPT5.6-Luna, with five attempts per task. The best-performing configuration in the six-model panel, GPT5.6-Luna with ReAct-3, passes 62.3% of attempts and solves 80.5% of tasks at least once (coverage) and 43.8% in all five attempts (consistency). ReAct-3 exceeds the one-pass Plan-and-Execute by 27.0 pass@1 points at 2.46 times the tokens. An expert audit of one sampled verifier-passing layout per covered task, across five ReAct-3 configurations, accepts 56.4% to 63.5% of them. The benchmark evaluates geometric and structural FET layout construction.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑