AI 中文总结
研究针对多模态大语言模型(MLLM)在工程多模态推理基准上的表现,构建MMArch基准,发现现有MLLM与人类专家存在差距,为相关研究提供评估基准。
AI 中文摘要
多模态大语言模型(MLLM)在工程图像上表现出色,但现有基准大多测试绘图识别、信息提取或合规性检查,未验证模型能否结合分布式视觉证据与工程原理得出结论。我们推出MMArch——一个覆盖建筑与土木工程十个子领域的基准,完全由同行评审论文中的图表构建。其1212个简答题项通过解耦的规划器-生成器流水线生成,并经自动筛选、盲对抗审计和专家审核验证,因此回答需感知相关证据、识别主导原理并应用,而非利用文本或单图捷径。我们将18种开源权重和专有MLLM与领域专家小组对比,发现存在巨大差距:最强开源模型准确率约30%,最佳专有系统为52%,而人类专家达95%,领先超40个百分点。错误分析显示,失败集中在原理应用和跨图表证据结合,而非证据定位,表明未来研究仍有较大空间。代码和数据可在此httpsURL获取。
英文摘要
Multimodal large language models (MLLMs) perform strongly on engineering imagery, yet existing benchmarks mostly test drawing recognition, information extraction, or compliance checking, leaving open whether models can combine distributed visual evidence with engineering principles to reach a conclusion. We introduce MMArch, a benchmark for architecture and civil engineering spanning ten subdomains and built entirely from figures in peer-reviewed papers. Its $1{,}212$ short-answer items are produced by a decoupled planner--writer pipeline and validated through automated screening, a blind adversarial audit, and expert review, so that answering requires perceiving the relevant evidence, identifying the governing principle, and applying it, not exploiting textual or single-figure shortcuts. Evaluating $18$ open-weight and proprietary MLLMs against a domain-expert panel, we find a wide gap: the strongest open-source model attains about $30\%$ and the best proprietary system $52\%$, while human experts reach $95\%$, more than forty points ahead. Our error analysis shows that failures concentrate in applying principles and combining evidence across figures rather than in locating it, pointing to substantial headroom for future research. Code and data are available at https://dcx-swjtu.github.io/MMArch/.