破坏模型以测试评判器:面向领域类图语义评估器的变异测试方法
Breaking Models to Test the Judge: A Mutation Testing Approach for Semantic Evaluators of Domain Class Diagrams
AI总结:
该研究针对领域类图语义评估器,提出一种变异测试方法,通过注入语义缺陷生成变体并评估评判器检测缺陷的能力,其结果与人工评估基本一致,可作为语义评判器分析的可扩展替代方案。
AI中文摘要:
在软件工程中,许多语义建模任务缺乏唯一的基准,因为人工评判既昂贵又主观。本文探索将变异测试作为评估模型语义评判器(例如基于大语言模型LLM的评判器)的可扩展替代方案。我们提出一种变异测试方法,该方法将受控语义缺陷注入领域类图中。从PlantUML类图与文本系统描述的配对开始,我们应用变异算子(例如移除一个类)生成有缺陷的变体。随后根据候选评判器检测注入缺陷的能力对其进行评估。我们为领域类图与文本描述的比较任务定义了11个变异算子,并针对传统的人工评判有效性评估对所提方法进行评估。在6种评判器配置(3种LLM与2种提示变体)中,自动化变异测试方法在识别性能更优的配置方面与人工评判基本一致。结果表明,变异测试可作为分析语义评判器的可扩展代理。
英文摘要:
In software engineering, many semantic modeling tasks lack a unique ground truth, as human judgments are both costly and subjective. This paper explores mutation testing as a scalable alternative for evaluating semantic judges (e.g., LLM-based) of models. We propose a mutation testing approach in which controlled semantic defects are injected into domain class diagrams. Starting from pairs of PlantUML class diagrams and textual system descriptions, we apply mutation operators (e.g., removing a class) to generate faulty variants. A candidate judge is then evaluated based on its ability to detect the injected defects. We define 11 mutation operators for the task of comparing a domain class diagram against a textual description and evaluate the proposed approach against a conventional manual assessment of judgment validity. Across six judge configurations (three LLMs and two prompt variants), the automated mutation testing approach is largely consistent with the manual assessment in identifying the better-performing configurations. The results suggest that mutation testing may serve as a scalable proxy for analyzing semantic judges.