arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.25516cs.SE

Swift 中不稳定测试的词汇特征

The Vocabulary of Flaky Tests in Swift

João Medeiros, Denini Silva, Breno Miranda

首次发表
浏览论文内容

中文总结 AI 辅助

本研究首次针对 Swift 语言评估基于词汇的机器学习不稳定测试预测,通过 15 个开源项目数据训练分类器,随机森林达到最优性能(F1=0.86),并揭示并发与同步词汇是主要不稳定信号,同时指出词汇预测的固有局限。

中文摘要 AI 辅助

不稳定测试在不修改代码的情况下产生非确定性结果,侵蚀了持续集成的信心并延迟交付。虽然基于词汇的机器学习预测已被证明对 Java 和 JavaScript 有效,但尚无研究针对 Swift 进行评估,Swift 的测试风格以 UI 和异步代码为主。我们通过重新执行和提交历史挖掘,从 15 个开源 Swift 项目中收集了 91 个不稳定测试和 22,349 个稳定测试,然后在分层 5 折交叉验证下,使用 TF-IDF 一元词+二元词特征训练了五个分类器(随机森林、决策树、朴素贝叶斯、SVM、KNN)。随机森林取得了最佳性能(精确率 = 0.92,F1 = 0.86,AUC = 0.95),并显著优于平凡基线,其中包括应用于最具信息量词汇的词汇阈值规则,确认了真实的判别信号(MCC = 0.75,而最佳基线为 0.08)。信息增益分析揭示了两种互补的信号类型。不稳定标记主要出现在不稳定测试中,包括并发原语(async、await)、基于期望的同步(expectation、fulfill)、错误传播(throws)和显式时间依赖(timeout、wait、now)。稳定标记主要是纯同步测试的断言词汇(xctassertequal),被视为不稳定的反证。错误分析表明,当不稳定性隐藏在测试体之外的共享基础设施中,或异步构造在确定性上下文中使用时,模型会失败,这暴露了词汇预测的内在局限性。这些结果将基于词汇的不稳定检测扩展到 Swift 生态系统,并刻画了其有效性和边界。

英文摘要

Flaky tests produce non-deterministic outcomes without code change, eroding CI confidence and delaying deliveries. While vocabulary-based machine learning prediction has proven effective for Java and JavaScript, no study has evaluated it for Swift, a language whose testing style is dominated by UI and asynchronous code. We collect 91 flaky and 22,349 stable tests from 15 open-source Swift projects via re-execution and commit-history mining, then train five classifiers (Random Forest, Decision Tree, Naive Bayes, SVM, KNN) on TF-IDF unigram+bigram features under stratified 5-fold cross-validation. Random Forest achieves the best performance (Precision = 0.92, F1 = 0.86, AUC = 0.95) and substantially outperforms trivial baselines, among them a vocabulary-threshold rule applied to the most informative tokens, confirming a genuine discriminative signal (MCC = 0.75 vs. 0.08 for the best baseline). Information-gain analysis reveals two complementary signal types. Flakiness markers appear predominantly in unstable tests and comprise concurrency primitives (async, await), expectation-based synchronisation (expectation, fulfill), error propagation (throws), and explicit timing dependence (timeout, wait, now). Stability markers, chiefly the assertion vocabulary of plainly synchronous tests (xctassertequal), count as evidence against flakiness. Error analysis shows that the model fails when flakiness is hidden in shared infrastructure outside the test body or when async constructs are used in a deterministic context, exposing the intrinsic limit of lexical prediction. These results extend vocabulary-based flakiness detection to the Swift ecosystem and characterise both its effectiveness and its boundaries.

发表机构

  • Centro de Informática Universidade Federal de Pernambuco(伯南布哥联邦大学信息学院)
  • Universidade Federal Rural de Pernambuco(伯南布哥联邦农村大学)

机构由 AI 辅助整理,请以论文原文为准。

↑