arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.02304cs.SEcs.AIcs.LG

SimuVerity:面向工程级Simulink模型生成的智能体基准测试

SimuVerity: Benchmarking Agents for Engineering-Grade Simulink Model Generation

  • Xi’an Jiaotong University(西安交通大学)

机构由 AI 辅助整理,请以论文原文为准。

Ruiqi Zhang, Jiahao Wang, Mingxuan Li, Haichen Luo, Chaoting Wang, Guoyu Mou, Keyu Lai, Hanchao Lv, Jiaxu Wang, Yibo Zheng, Aijun Yang, Xiaohua Wang

AI总结:

SimuVerity提出101个跨十领域的Simulink模型生成任务及分层评估器,测试六个智能体系统,最佳得分仅42.86,揭示结构相似性不能代表工程性能,并诊断出能力瓶颈与视觉布局问题。

AI中文摘要:

现有的Simulink基准测试主要评估生成的模型能否编译、执行或与参考模型相似。这些标准并不能确定模型是否满足其工程需求。我们引入了SimuVerity,这是一个包含101个跨十个工程领域的文本到可执行Simulink模型生成任务的基准测试。对于每个任务,可执行系统配置文件为工程规范和四类原生仿真场景提供了基础。一个分层评估器首先检查工件交付、原生可执行性和工程资格,然后从六个维度对合格模型进行评分,涵盖准确性、输出质量、机制保真度、控制与因果完整性、操作域鲁棒性和动态响应。我们使用SimuVerity评估了六个智能体系统。最佳系统的总体得分仅为42.86。结果表明,结构相似性是工程性能的一个较差替代指标:能力瓶颈既出现在生成合格实现方面,也出现在合格后满足多维需求方面。同时,一些高得分模型仍表现出严重的视觉布局混乱。SimuVerity为评估智能体的工程能力和诊断可执行Simulink模型生成中的故障提供了系统性基础。

英文摘要:

Existing Simulink benchmarks mainly evaluate whether generated models compile, execute, or resemble a reference model. These criteria do not establish whether a model satisfies its engineering requirements. We introduce SimuVerity, a benchmark of 101 text-to-executable Simulink model-generation tasks across ten engineering domains. For each task, executable-system profiles ground the engineering specification and four families of native simulation scenarios. A hierarchical evaluator first checks artifact delivery, native executability, and engineering qualification, then scores qualified models across six dimensions covering accuracy, output quality, mechanistic fidelity, control and causal integrity, operating-domain robustness, and dynamic response. We evaluate six agent systems with SimuVerity. The best system achieves an overall score of only 42.86. The results show that structural similarity is a poor proxy for engineering performance: capability bottlenecks arise both in producing qualified implementations and in satisfying multidimensional requirements after qualification. Meanwhile, some high-scoring models still exhibit severe visual-layout disorder. SimuVerity provides a systematic basis for assessing agents' engineering capabilities and diagnosing failures in executable Simulink model generation.

补充信息

↑