arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

CircuitReason-1k:电路中的长程视觉到符号推理基准测试

CircuitReason-1k: Benchmarking Long-Horizon Visual-to-Symbolic Reasoning inElectrical Circuits

Xinqi Yang, Kang An, Tengyue Wang, Zhongyu Yang, Chenxu Du, Yuanchi Zhu, Hebao Zhu, Ziliang Wang, Faqiang Qian, Yunli Yang, Qibing Ren

arXiv 2608.09374首次发表:更新:

AI 中文总结

CircuitReason-1k是含1000个真实教材电路问题的基准,用于评估多模态模型的长程视觉到符号推理,现有模型在长程问题上表现差,该基准为相关测试提供了聚焦平台。

AI 中文摘要

电路分析仅识别图像中的组件是不够的,求解器必须将符号和标签与实际对应,恢复潜在拓扑结构,选择物理模型,构建耦合方程,传播中间量,并保留单位、符号、方向和相位约定。我们推出CircuitReason-1k,这是一个包含1000个真实教材问题的基准,用于评估这一完整的长程视觉到符号推理过程。每个问题搭配一个或多个电路图,以及一个独立的问题、一个按类型或语义指定的答案和一个参考解题过程。以证据为核心的构建流程将问题、图表和解题过程对齐,而面向推理的分类法则按电路类型和依赖深度对问题进行组织。评估结合了保守的类型评分与身份盲多模型语义共识,分母保留所有问题。在三个商业聊天机器人系统和六个开源多模态大语言模型中,得分最高的系统达到84.8%的准确率。然而,长程问题的性能持续下降,定性分析显示在拓扑到目标的绑定、物理约定以及后期输出传播方面存在持续的失败。CircuitReason-1k为衡量多模态模型能否将技术视觉证据转化为持续的、物理上有效的符号推理提供了一个聚焦的测试平台。代码可在GitHub - CircuitReason/CircuitReason1K获取。

英文摘要

Electrical circuit analysis requires more than recognizing components in an image. A solver must ground symbols and labels, recover latent topology, select a physical model, formulate coupled equations, propagate intermediate quantities, and preserve units, signs, directions, and phase conventions. We introduce \benchmark, a benchmark of 1,000 authentic textbook problems for evaluating this complete long-horizon visual-to-symbolic reasoning process. Each problem pairs one or more circuit diagrams with a self-contained question, a typed or semantically specified answer, and a reference worked solution. An evidence-first construction pipeline aligns questions, figures, and solutions, while a reasoning-oriented taxonomy organizes problems by circuit type and dependency depth. Evaluation combines conservative typed scoring with identity-blinded multi-model semantic consensus, retaining every problem in the denominator. Across three commercial chatbot systems and six open-source multimodal large language models, the highest-scoring system reaches 84.8\% accuracy. However, performance consistently deteriorates on long-horizon problems, and qualitative analysis exposes persistent failures in topology-to-target binding, physical conventions, and late-stage output propagation. \benchmark{} provides a focused testbed for measuring whether multimodal models can transform technical visual evidence into sustained, physically valid symbolic reasoning. Code are available at GitHub - CircuitReason/CircuitReason1K.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑