arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.25528cs.SE

评估 Shaker 在 Python 项目中的不稳定测试检测能力

Evaluating Shaker for Flaky Test Detection in Python Projects

Gabriela Leal, Denini Silva, Leopoldo Teixeira

首次发表
浏览论文内容

中文总结 AI 辅助

首次实证评估 Shaker 在 Python 中的不稳定测试检测,发现其相对普通重执行无显著优势,并揭示跨环境重用真实数据集会降低工具召回率的风险。

中文摘要 AI 辅助

不稳定测试在代码未变更的情况下会非确定性地通过或失败,从而削弱对测试套件的信任,并增加每次失败的成本。Shaker 通过注入资源争用(CPU、内存和 I/O 压力)来放大由并发执行引起的非确定性,从而检测不稳定测试,据报告,在 Java 和 Android 基准测试中,它能检测出 95% 的不稳定测试,而普通重新执行(ReRun)仅为 37.5%。我们首次对 Shaker 在 Python 上的应用进行了实证评估。从 Gruber 等人的真实数据集中抽取非顺序依赖的不稳定测试,我们在配对设计中比较了 Shaker 与预算匹配的 ReRun 基线,使两种技术获得相同数量的测试执行:每个测试在每种技术下各运行 100 次,共 137 个测试。按照为 Java 和 Android 配置的方式,Shaker 相对于普通重新执行没有提供统计上显著的检测优势(37.2% 对比 35.8%;McNemar 精确检验 p = 0.84)。两个发现解释了原因。首先,在独立硬件上,无论采用哪种技术,真实数据集中不到一半的不稳定测试能够复现为不稳定,且大多数无法复现的测试在 100 次运行中从未出现分歧。其次,能够复现的测试主要由网络交互和随机性引起的不稳定性主导,而非 Shaker 所针对的并发问题。除了工具本身,这还暴露了该领域的一个更广泛的风险:在不同执行环境中重用不稳定测试的真实数据集,会悄然将真正的不稳定测试转化为表面上的真阴性,从而降低任何工具测得的召回率。

英文摘要

Flaky tests pass or fail non-deterministically on unchanged code, eroding trust in test suites and inflating the cost of every failure. Shaker detects them by injecting resource contention (CPU, memory, and I/O stress) to amplify non-determinism caused by concurrent execution, and was reported to detect 95% of the flaky tests in a Java and Android benchmark against 37.5% for plain re-execution (ReRun). We present the first empirical evaluation of Shaker for Python. Drawing non-order-dependent flaky tests from the ground-truth dataset of Gruber et al., we compare Shaker against a budget-matched ReRun baseline in a paired design, giving both techniques the same number of test executions: Each of 137 tests is run 100 times under each. As configured for Java and Android, Shaker provides no statistically significant detection advantage over plain re-execution (37.2% vs. 35.8%; McNemar exact p = 0.84). Two findings explain why. First, fewer than half of the ground-truth flaky tests reproduce as flaky at all on independent hardware under either technique, and most of the tests that fail to reproduce never diverge once across 100 runs. Second, the tests that do reproduce are dominated by flakiness from network interactions and randomness rather than the concurrency Shaker targets. Beyond the tool, this exposes a broader hazard for the field: reusing a flaky-test ground truth across execution environments silently converts genuine flaky tests into apparent true negatives, deflating any tool's measured recall.

发表机构

  • Centro de Informática Universidade Federal de Pernambuco(伯南布哥联邦大学计算机中心)
  • Universidade Federal Rural de Pernambuco(伯南布哥联邦农村大学)

机构由 AI 辅助整理,请以论文原文为准。

↑