AI 中文总结
ModBench是用于构建Modelica基准数据集的流水线,经其处理Modelica标准库得到含85562个类快照的公开数据集,可支持模型演化分析等相关研究。
AI 中文摘要
基于方程的网络物理系统建模语言(如Modelica)的研究受限于缺乏经过整理的基准数据集,这限制了对模型演化与发展的实证洞察。我们提出ModBench,一种从Modelica库的Git仓库中挖掘以生成模型快照基准数据集的流水线。该流水线包含三个步骤:(1)过滤仓库提交记录,保留人工编写且与Modelica相关的修订;(2)提取符合仿真要求的类;(3)构建Modelica类的规范表示。为进行实证验证,我们将ModBench应用于Modelica标准库(MSL),并报告生成的数据集:其涵盖自Modelica语言v3以来的完整提交历史,包含85562个不同的类快照,且带有可追溯至原始模型和Git元数据的链接。该数据集、其API及数据生成都流水线已公开,以支持未来关于模型演化分析、编译器测试以及模型自动修复或生成的研究。
英文摘要
Research on equation-based cyber-physical systems modeling languages, such as Modelica, is constrained by the lack of curated benchmark datasets. This limits empirical insight into the evolution and development of models. We address this gap with ModBench, a pipeline that mines Git repositories of Modelica libraries to produce benchmark datasets of model snapshots. The pipeline (1) filters repository commits to retain human-authored, Modelica-relevant revisions; (2) extracts simulation-eligible classes; and (3) builds canonical representations of Modelica classes. For empirical validation, we applied ModBench to the Modelica Standard Library (MSL) and report the resulting dataset, spanning the full commit history (since Modelica language v3), with 85,562 distinct class snapshots, and links enabling traceability to original models and Git metadata. The dataset, its API, and the data generation pipeline are publicly available to support future research on model evolution analysis, compiler testing, and automated model repair or generation.
CommentsExtended abstract accepted at SAM 2026, co-located with MODELS 2026. To appear in the ACM/IEEE MODELS 2026 Companion Proceedings