arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2509.22339cs.CV

CircuitSense: 一种层次化的MLLM基准,连接视觉理解和符号推理在工程设计过程

CircuitSense: A Hierarchical MLLM Benchmark Bridging Visual Comprehension and Symbolic Reasoning in Engineering Design Process

  • Northeastern University(东北大学)
  • Brookhaven National Laboratory(布鲁克海文国家实验室)

机构由 AI 辅助整理,请以论文原文为准。

Arman Akbari, Jian Gao, Yifei Zou, Mei Yang, Jinru Duan, Dmitrii Torbunov, Yanzhi Wang, Yihui Ren, Xuan Zhang

更新

AI总结:

CircuitSense提出一个层次化基准,评估工程设计中视觉理解与符号推理的能力,揭示闭源模型在符号推导上的不足,强调数学理解对电路综合的重要性。

AI中文摘要:

工程设计通过从系统规范到组件实现的层次化抽象进行操作,需要在每个层次上结合视觉理解和数学推理。尽管多模态大语言模型(MLLMs)在自然图像任务上表现出色,但它们从技术图纸中提取数学模型的能力仍未被探索。我们提出了CircuitSense,一个全面的基准,通过8060多个问题评估电路理解的整个层次,从组件级别的原理图到系统级别的方框图。我们的基准独特地考察了完整的工程工作流程:感知、分析和设计,特别强调从视觉输入中推导出符号方程这一关键但未被充分探索的能力。我们介绍了一种分层的合成生成管道,包括基于网格的原理图生成器和方框图生成器,后者具有自动推导的符号方程标签。对六个最先进的MLLMs的全面评估,包括闭源和开源模型,揭示了视觉到数学推理的根本限制。闭源模型在涉及组件识别和拓扑识别的感知任务中达到超过85%的准确率,但其在符号推导和分析推理上的表现低于19%,暴露了视觉解析和符号推理之间的关键差距。具有更强符号推理能力的模型在设计任务中的准确性更高,证实了数学理解在电路综合中的基本作用,并将符号推理确立为工程能力的关键度量标准。

英文摘要:

Engineering design operates through hierarchical abstraction from system specifications to component implementations, requiring visual understanding coupled with mathematical reasoning at each level. While Multi-modal Large Language Models (MLLMs) excel at natural image tasks, their ability to extract mathematical models from technical diagrams remains unexplored. We present \textbf{CircuitSense}, a comprehensive benchmark evaluating circuit understanding across this hierarchy through 8,006+ problems spanning component-level schematics to system-level block diagrams. Our benchmark uniquely examines the complete engineering workflow: Perception, Analysis, and Design, with a particular emphasis on the critical but underexplored capability of deriving symbolic equations from visual inputs. We introduce a hierarchical synthetic generation pipeline consisting of a grid-based schematic generator and a block diagram generator with auto-derived symbolic equation labels. Comprehensive evaluation of six state-of-the-art MLLMs, including both closed-source and open-source models, reveals fundamental limitations in visual-to-mathematical reasoning. Closed-source models achieve over 85\% accuracy on perception tasks involving component recognition and topology identification, yet their performance on symbolic derivation and analytical reasoning falls below 19\%, exposing a critical gap between visual parsing and symbolic reasoning. Models with stronger symbolic reasoning capabilities consistently achieve higher design task accuracy, confirming the fundamental role of mathematical understanding in circuit synthesis and establishing symbolic reasoning as the key metric for engineering competence.

↑