发表机构
University of Innsbruck; University of Augsburg(因斯布鲁克大学; 奥格斯堡大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究构建了全自动基准MecEng,评估32个开放权重及2个专有LLMs的机械工程认知,发现其刚体任务成功率达86.0%,但柔性多体任务仍具挑战,且该能力正快速提升但仍易出错。
AI 中文摘要
大语言模型(LLMs)在成熟的代码生成和数学推理基准测试中表现良好,但它们在力学和空间几何方面的能力(本文称之为“机械工程认知”)尚未被系统量化。本文提出MecEng,这是一个全自动基准,用于评估LLMs从参数化文本描述创建多体仿真模型的能力。该基准包含84个通用任务,分为三个难度级别,涵盖从带关节和接触的刚体系统到需要精确3D几何生成、四面体有限元网格划分以及机械零件的Hurty-Craig-Bampton模型降阶的柔性多体系统。一个专用流程利用LLMs通过Netgen从文本生成可用于仿真的几何结构,并为Exudyn代码构建多体系统模型,随后在多个层面上对照专家基准进行验证:包括图节点注释的系统图同构、数值解以及质量、几何和固有频率等零件特定指标。本文总共评估了32个开放权重LLMs和2个专有LLMs。在刚体任务中,最佳开放权重模型的总体成功率为86.0%,而最强的专有模型为91.4%,不过柔性多体任务仍然难度大得多。额外研究量化了采样温度、推理、提示设计、模型规模和LLMs发布日期的影响。结果表明,当前LLMs的机械工程认知能力正在快速提升,但仍然容易出错。
英文摘要
Large Language Models (LLMs) perform well on established code-generation and mathematical-reasoning benchmarks, but their capabilities in mechanics and spatial geometry, here denoted as mechanical engineering awareness, has not been quantified systematically. We present MecEng, a fully automated benchmark that evaluates LLMs on the creation of multibody simulation models from parameterized textual descriptions. The benchmark comprises 84 generic tasks on three difficulty levels, ranging from rigid-body systems with joints and contact to flexible multibody systems that require exact 3D geometry generation, tetrahedral finite-element meshing, and Hurty-Craig-Bampton model order reduction of machine parts. A dedicated pipeline with LLMs generates simulation-ready geometry from text using Netgen, and builds multibody system models for the code Exudyn, which are then verified against expert ground truth on several levels: system-graph isomorphism including graph node annotations, numerical solutions, and part-specific measures such as mass, geometry, and eigenfrequencies. In total, 32 open-weight and two proprietary LLMs are evaluated. On rigid-body tasks, the best open-weight model obtains an overall success rate of 86.0%, compared to 91.4% for the strongest proprietary model, while flexible multibody tasks remain considerably harder. Additional studies quantify the influence of sampling temperature, reasoning, prompt design, model size, and LLM-release date. The results indicate rapidly improving, but still error-prone, mechanical engineering awareness of current LLMs.