发表机构
School of Computer Science and Engineering, Macau University of Science and Technology; Macau University of Science and Technology Zhuhai MUST Science and Technology Research Institute; School of Information Science and Technology, Dalian Maritime University; School of Science, University of Tokyo; Department of Electrical and Computer Engineering, University of Alberta(澳门科技大学计算机科学与工程学院; 澳门科技大学珠海澳门科技大学研究院; 大连海事大学信息科学与技术学院; 东京大学理学部; 阿尔伯塔大学电气与计算机工程系)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究系统综述组合交互测试的黑盒评估指标,并通过涵盖32个测试场景和295,624个测试套件的实证分析,发现值组合覆盖率(VCC)是有效的预测指标,且基于分布的指标计算成本更低,最终为从业者提供了指标选择指南。
AI 中文摘要
组合交互测试(CIT)是一种黑盒测试方法,近年来在研究和实践中均受到广泛关注。其主要目标是构建有效的组合测试套件,以检测由参数交互引起的软件故障。作为CIT测试过程的基本组成部分,评估指标在评估和比较组合测试套件以及评价各种测试生成技术方面发挥着关键作用。对于CIT从业者而言,鉴于可用选项的多样性,选择合适的评估指标既重要又具有挑战性。然而,此前尚无工作系统性地解决这一问题。为填补这一空白,本文首先对组合测试套件的黑盒评估指标进行了全面综述,提供了严格的定义、清晰的分类、说明性示例和复杂度分析。随后,我们开展了一项涉及八个开源项目、涵盖32个测试场景和295,624个组合测试套件的广泛实证研究。在该研究中,我们使用两种相关性度量考察了每个静态评估指标与故障检测有效性之间的相关性。实验结果表明,值组合覆盖率(VCC)指标可作为测试套件评估的有效预测因子。然而,基于分布的指标通常比基于交互覆盖率的指标具有更低的计算成本。合适指标的选择还应考虑测试套件的固有属性,因为这些特征可能显著影响评估的有效性。最后,我们提供了实用指南,以帮助CIT从业者选择合适的评估指标来评估或比较组合测试套件。
英文摘要
Combinatorial interaction testing (CIT) is a black-box testing method that has received extensive attention in both research and practice over recent years. Its primary objective is to construct an effective combinatorial test suite that detects software failures caused by parameter interactions. As a fundamental component of the CIT testing process, the evaluation metric plays a critical role in assessing and comparing combinatorial test suites, as well as in evaluating various test generation techniques. For CIT practitioners, selecting an appropriate evaluation metric is both important and challenging, given the wide variety of available options. Nevertheless, no prior work has systematically addressed this problem. To fill this gap, this paper first provides a comprehensive survey of black-box evaluation metrics for combinatorial test suites, offering rigorous definitions, clear classifications, illustrative examples, and complexity analyses. We then conduct an extensive empirical study involving eight open-source projects, encompassing 32 test scenarios and 295,624 combinatorial test suites. In this study, we examine the correlation between each static evaluation metric and fault-detection effectiveness using two correlation measures. Experimental results show that the Value Combination Coverage (VCC) metric serves as a valid predictor for test-suite evaluation. However, distribution-based metrics generally incur lower computational costs than interaction coverage-based ones. The choice of an appropriate metric should also account for the test suite's inherent properties, as these characteristics can substantially influence the effectiveness of the evaluation. Finally, we provide practical guidelines to assist CIT practitioners in selecting suitable evaluation metrics for assessing or comparing combinatorial test suites.