基于大语言模型的流体系统仿真代码生成:模型与提示策略的基准测试
Simulation Code Generation for Fluid Systems using Large Language Models: Benchmarking Models and Prompting Strategies
查看机构详情
- German Aerospace Centre (DLR)(德国航空航天中心)
- Technische Hochschule Würzburg-Schweinfurt(维尔茨堡-施韦因富特应用技术大学)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
本研究通过系统比较10种大语言模型与6种提示策略,评估其将流体系统模型图表示转换为WNTR、Modelica代码的能力,为相关设计流程提供指导,同时指出仿真保真度仍存差距。
中文摘要 AI 辅助
大语言模型(LLMs)已展现出从自然语言规范生成语法正确代码的强大能力。本研究探索如何利用LLMs将流体系统模型的中性图表示自动转换为两种广泛采用的仿真环境的可执行代码:Python库WNTR和Modelica标准库。我们对10种最先进的LLMs和6种提示策略进行了系统比较,这些策略在提供的上下文信息(如代码或文档)上存在差异。针对每种配置,我们使用一套软件质量指标评估生成的代码,并通过复现基准流体系统场景验证所得仿真模型的功能保真度。研究结果为希望将LLM驱动的代码合成集成到基于模型的设计流程中的研究人员和工程师提供了具体指导。尽管表现最佳的配置达到了可接受的语法质量,但我们发现仿真保真度仍存在显著差距。
英文摘要
Large language models (LLMs) have demonstrated a strong ability to generate syntactically correct code from natural-language specifications. In this study, we explore how LLMs can be harnessed to automatically translate a neutral graph representation of fluid system models into executable code for two widely adopted simulation environments: the Python library WNTR and the Modelica Standard Library. We conduct a systematic comparison of ten state-of-the-art LLMs and six prompting strategies that differ in the contextual information supplied (e.g., code or documentation). For each configuration we assess the generated code using a suite of software-quality metrics and we validate the functional fidelity of the resulting simulation models by reproducing benchmark fluid system scenarios. Our findings offer concrete guidance for researchers and engineers seeking to integrate LLM-driven code synthesis into model-based design pipelines. While the best-performing configurations achieve acceptable syntactic quality, we observe substantial gaps remain in simulation fidelity.