arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.03078cs.CV

LDU-Bench:面向不同布局电路背景下光刻缺陷理解的多模态大语言模型评估基准

LDU-Bench: Multimodal LLM Evaluation for Lithography Defect Understanding under Layout-Varying Circuit Backgrounds

Huanglong Ji, Botong Zhao, Shujing Lv, Yue Lv

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出LDU-Bench这一多模态基准,将光刻审查工作流程分解为四项任务,评估发现现有多模态大语言模型的缺陷分类能力无法稳定迁移至下游审查阶段,为工业MLLM评估提供了统一平台。

中文摘要 AI 辅助

多模态大语言模型(Multimodal large language models)在工业异常检测中展现出强大的缺陷识别能力。然而在光刻审查中,仅判断图像是否存在缺陷不足以满足工程检测需求,模型还需理解缺陷的形态、空间位置以及由可见证据支持的潜在成因。为此,本文提出LDU-Bench,这是一个面向光刻缺陷理解的多任务多模态基准。LDU-Bench由真实光刻图像和集成电路审查图像构建而成,将审查工作流程分解为四个独立任务:缺陷分类(defect triage)、形态识别(morphology recognition)、粗定位(coarse localization)以及图像条件成因分析(image-conditioned cause analysis)。该基准采用任务级指标、诊断读数以及光刻闭合分数(Lithography Closure Score, LCS)对模型进行系统性评估。实验结果表明,尽管现有多模态大语言模型(MLLMs)能够较为可靠地执行缺陷分类任务,但该能力无法稳定迁移至下游审查阶段,形态对齐、有效定位以及证据到成因的映射仍是主要瓶颈。进一步诊断显示,该能力缺口并非单一指标的波动,而是反映了跨语义层次的结构化理解不足。总体而言,LDU-Bench为评估工业多模态大语言模型在光刻审查链中的可用性、失效点及能力边界提供了可量化且具备诊断性的统一平台。

英文摘要

Multimodal large language models have demonstrated strong defect recognition capability in industrial anomaly detection. However, in lithography review, merely determining whether an image contains a defect is insufficient for engineering inspection; models must also understand defect morphology, spatial location, and the potential causes supported by visible evidence. To this end, this paper proposes LDU-Bench, a multi-task multimodal benchmark for lithography defect understanding. Constructed from real lithography and integrated-circuit review images, LDU-Bench decomposes the review workflow into four independent tasks: defect triage, morphology recognition, coarse localization, and image-conditioned cause analysis. It systematically evaluates models using task-level metrics, diagnostic readouts, and the Lithography Closure Score (LCS). Experimental results show that although existing MLLMs can perform defect triage relatively reliably, this ability does not stably transfer to downstream review stages. Morphology alignment, effective localization, and evidence-to-cause mapping remain the major bottlenecks. Further diagnostics indicate that this capability break is not a fluctuation of a single metric, but reflects insufficient structured understanding across semantic levels. Overall, LDU-Bench provides a quantifiable and diagnostic unified platform for evaluating the usability, failure points, and capability boundaries of industrial MLLMs in lithography review chains.

补充信息

↑