arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

GReFEM:多模态大语言模型作为用于物理引导的3D网格细化的零样本语义助手

GReFEM: Multimodal LLMs as Zero-Shot Semantic Assistants for Physics-Guided 3D Mesh Refinement

Kartik Bali, Mahish K. Guru, Christian J Cyron, Roland Aydin

arXiv 2607.08798首次发表:更新:

发表机构

Institute of Material Systems Modeling; Helmholtz Zentrum Hereon; Institute of Material and Process Design; German Research Center for Artificial Intelligence(材料系统建模研究所; 海德堡研究中心Hereon; 材料与加工设计研究所; 德国人工智能研究中心)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究探索多模态大语言模型能否作为有限元网格细化的零样本几何代理,引入GReFEM框架及orthoViews视图选择模块,经实证评估发现其展示出强大零样本能力,能准确遵循复杂指令,比盲目启发式方法更精确,定义了基础模型在自动模拟中作为语义助手的前沿。

AI 中文摘要

自适应体积有限元网格划分是计算机辅助工程与分析中的关键步骤,决定给定问题的计算预算。传统上需要迭代偏微分方程求解器或在大规模模拟数据上训练的强监督、数据驱动的替代模型。多模态大语言模型在二维视觉任务中表现出色,但基于几何理解和物理进行语义定位区域的零样本能力仍是开放问题。本研究探索现成的多模态大语言模型的高级语义理解能否作为有限元网格细化的可行零样本几何代理。为此引入GReFEM框架,利用多模态大语言模型基于物理引导的文本提示在视觉上定位应力关键区域。还引入orthoViews视图选择模块弥合二维多模态大语言模型预训练与三维几何之间的差距。通过对不同CAD几何、加载情况和最先进的多模态大语言模型进行深入实证评估,并与在严格匹配细化预算下调整后的几何启发式方法比较。结果表明多模态大语言模型展示出强大的零样本能力,能准确遵循复杂的空间物理指令,比盲目启发式方法更精确地分离与应力相关的特征。本研究通过映射多模态大语言模型在物理基础方面的成功与当前局限,定义了基础模型在自动模拟工作流程中作为语义助手的前沿领域。

英文摘要

Adaptive volumetric finite element meshing is a critical step in computer-aided engineering and analysis that dictates the computational budget of a given problem. It traditionally requires iterative PDE solvers or heavily supervised, data-driven surrogates trained on large-scale simulation data. While Multimodal Large Language Models (MLLMs) excel in 2D visual tasks, their zero-shot capability to semantically ground regions based on geometric understanding and physics remains an open question. Overall, this study explores a significant question: can the high-level semantic understanding of off-the-shelf MLLMs serve as a viable, zero-shot geometric proxy for finite element mesh refinement? To investigate this, we introduce GReFEM (Geometric Reasoning Enhanced Multimodal LLMs for Finite Element Meshing), a framework that utilizes MLLMs to visually localize stress-critical regions based on physics-guided textual prompts. To bridge the gap between 2D MLLM pre-training and 3D geometries, we introduce orthoViews, a view-selection module that maximizes the observability of key geometric features. We conduct an in-depth empirical evaluation across diverse CAD geometries, loading cases, and SOTA MLLMs, comparing them against a tuned geometric heuristic under a strict, matched refinement budget. Our findings reveal that MLLMs demonstrate robust zero-shot capacity to accurately follow complex spatial-physical instructions, isolating stress-relevant features with higher precision than blind heuristics. By mapping both the successes and current limitations of MLLMs in physical grounding, this study defines the frontier of foundation models as semantic assistants in automated simulation workflows.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑