PolyBridgeBench:面向物理基础桥梁设计的多模态大语言模型基准测试
PolyBridgeBench: Benchmarking Multimodal LLMs for Physics-Grounded Bridge Design
- Fudan University(复旦大学)
- Shanghai Innovation Institute(上海创新研究院)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
该研究提出PolyBridgeBench基准,用于评估多模态大语言模型在物理模拟中设计承重桥梁及失败后修复的能力,实验发现模型在动态成功与恢复方面存在明显不足。
AI中文摘要:
多模态大语言模型(MLLMs)在视觉理解和结构化生成方面表现出色,但这些能力并不能确定工程设计在执行时是否有效。现有基准测试评估空间推理、结构有效性或基于物理的构建,但无法确定MLLMs能否综合出完整的承重结构,并在模拟器执行暴露失败后进行修复。我们引入了PolyBridgeBench,一个用于多模态桥梁设计的可执行基准测试。模型接收视觉场景和结构化工程约束,生成完整的节点-构件-材料拓扑。确定性合法性检查在原生动态物理模拟中控制执行。执行失败后,基准测试返回失败滚动的时间视觉证据,并在固定交互预算下评估修复能力。对确定性有效性、动态功能成功和失败后恢复的独立测量,可识别设计失败的阶段。在189个关卡中对六个代表性MLLMs进行的实验,揭示了确定性有效性与动态成功之间的显著差距,对材料预算的明显敏感性,以及在主要严格预算设置下有限的失败后恢复能力。
英文摘要:
Multimodal large language models, or MLLMs, perform well at visual understanding and structured generation, yet these capabilities do not establish whether an engineering design will work when executed. Existing benchmarks assess spatial reasoning, structural validity, or physics-grounded construction, but they do not determine whether MLLMs can synthesize complete load-bearing structures and repair them after simulator execution exposes a failure. We introduce PolyBridgeBench, an executable benchmark for multimodal bridge design. A model receives a visual scene and structured engineering constraints and generates a complete node--member--material topology. Deterministic legality checks gate execution in a native dynamic physics simulation. Following an execution failure, the benchmark returns temporal visual evidence from the failed rollout and evaluates repair under a fixed interaction budget. Separate measurements of deterministic validity, dynamic functional success, and post-failure recovery identify the stage at which design fails. Experiments with six representative MLLMs across 189 levels expose a substantial gap between deterministic validity and dynamic success, pronounced sensitivity to material budgets, and limited post-failure recovery under the primary strict-budget setting.