arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.03718cs.CEcs.CLphysics.comp-ph

通用框架之外,CAE仿真智能体还需要什么?

What Do CAE Simulation Agents Really Need Beyond a Generic Harness?

Jiasheng Shi, Tianhan Zhang

首次发表
浏览论文内容

中文总结 AI 辅助

该研究探究CAE仿真智能体在通用框架外的需求,发现单智能体框架性能优于多智能体专用系统,执行反馈修复是关键,求解器教程类领域知识仍有显著增益。

中文摘要 AI 辅助

计算机辅助工程(CAE)仿真是工程领域规模最大、要求最高的领域之一,设置OpenFOAM、FEniCS或COMSOL等求解器需要真正的专业知识。大语言模型(LLM)智能体有望将自然语言请求转化为可运行的仿真,近期的CAE智能体加入了仿真专用机制:多智能体分解、领域检索和脚本化反思。这些机制适用于弱基础模型;现代框架已提供多轮推理、工具使用和执行反馈。本文探究CAE仿真智能体在通用框架之外仍需什么。在信息访问和修复预算固定的情况下,单智能体框架的性能与多智能体专用系统相当或更优(FoamBench为96.4%,多智能体系统为88.2%)。消融实验将此归因于框架已具备的能力:执行反馈修复使FoamBench从无修复轮次的71.8%提升至96.4%,而脚本化反思无增益。唯一仍有帮助的输入是作为求解器教程提供的领域知识,这是观测到的最大增益(从80.9%提升至96.4%)。

英文摘要

Computer-aided engineering (CAE) simulation is among the largest and most demanding areas of engineering, where setting up a solver such as OpenFOAM, FEniCS, or COMSOL takes real expertise. Large language model (LLM) agents promise to turn a natural-language request into a working simulation, and recent CAE agents add simulation-specific machinery: multi-agent decomposition, domain retrieval, and scripted reflection. That machinery suited weak base models; modern harnesses already supply multi-turn reasoning, tool use, and execution feedback. We ask what a CAE simulation agent still needs beyond a generic harness. With information access and repair budget held fixed, a single-agent harness matches or beats multi-agent specialized systems (FoamBench 96.4\% vs.\ 88.2\%). Ablations trace this to capabilities the harness already provides: execution-feedback repair lifts FoamBench from 71.8\% with no repair round to 96.4\%, while scripted reflection adds nothing. The one input that still helps is domain knowledge supplied as solver tutorials, our largest measured gain (80.9\% to 96.4\%).

发表机构

  • DP Technology(达梦科技)
  • School of Astronautics, Beihang University(北京航空航天大学宇航学院)
  • AI for Science Institute(人工智能科学研究院)

机构由 AI 辅助整理,请以论文原文为准。

↑