arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

面向自动化领域建模的标准化评估:引入一个基准

Towards Standardized Evaluation in Automated Domain Modeling: Introducing a Benchmark

Vasiliy Seibert

arXiv 2608.15255首次发表:更新:

发表机构

Institute for Software and Systems Engineering, TU Clausthal(克劳斯塔尔工业大学软件与系统工程研究所)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文针对自动化领域建模缺乏标准化基准的问题,引入整合两类数据集的基准,评估多种建模方法,为该领域研究提供可复用的研究成果。

AI 中文摘要

领域建模在领域驱动设计中发挥着关键作用,用于捕获特定领域内的核心实体及其关系。尽管自动化领域建模已取得进展,但缺乏标准化基准阻碍了现有方法的对比评估。本文引入一个基准以填补该空白:该基准整合了Zenodo上由Calamo、Mecella和Snoeck的Text2UML项目发布的含45条记录的Golden UML Modelset(Verbruggen等人,2025),以及Chen等人(2023a,b)的含8条记录的参考档案,可在不同复杂度和规模下评估自动化领域建模方法。给定自然语言描述,任务是生成对应领域模型,每条描述均提供参考领域模型作为真值,使用指标对比生成模型与真值模型。为展示基准的实用性,本文评估了多种自动化领域建模方法,包括基于启发式规则的方法和大语言模型(LLM)驱动的策略。该基准符合FAIR4RS建议(Chue Hong等人,2022),作为研究成果提供以鼓励复用,支持自动化领域建模的未来研究。

英文摘要

Domain modeling plays an essential role in domain-driven design, capturing essential entities and their relationships within a specific domain. Despite advancements in automated domain modeling, the absence of standardized benchmarks has hindered the comparative assessment of existing approaches. This paper introduces a benchmark designed to address this gap. The benchmark combines the 45-record Golden UML Modelset (Verbruggen et al., 2025) on Zenodo, as distributed by the Text2UML project of Calamo, Mecella, and Snoeck (Calamo et al., 2025), with the 8-record reference archive of Chen et al. (Chen et al., 2023a,b), enabling the evaluation of automated domain modeling approaches across different levels of complexity and scale. Given a natural language description, the task is to generate a corresponding domain model. For each description, a reference domain model is provided as ground truth. A metric is used to compare the generated domain model with the corresponding ground-truth model. To demonstrate the utility of the benchmark, we evaluate multiple automated domain modeling approaches, including heuristic rule-based methods and LLM-driven strategies. In accordance with the FAIR4RS recommendations (Chue Hong et al., 2022), the benchmark is provided as a research artifact to encourage reuse and support future research on automated domain modeling.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑