AI 中文总结
研究基于代码检测不稳定测试的局限性,分析发现常用基准测试有捷径、评估协议夸大通用性。整理C-IDoFT和CI日志数据集,重新构建不稳定预测,围绕观察到的失败及执行环境下测试失败可能性展开,数据集和挖掘方法支持此方向。
AI 中文摘要
不稳定测试在相同代码版本上通过和失败,削弱了测试结果信号并扰乱持续集成(CI)管道。基于代码的不稳定检测工具虽有不错的基准测试结果,但实际应用有限。我们认为该领域研究的问题有误,不稳定并非测试代码的静态属性。分析三种基于代码的检测工具后发现,常用基准测试存在使F1分数虚高的捷径,评估协议也夸大了通用性。我们精心整理了两个数据集。第一个是C-IDoFT(来自57个GitHub项目的54,468个单元测试),保留开发者确认的不稳定测试子集,仅从重复执行中重建非不稳定类。在已发布的交叉验证协议下,我们对两个已发布检测工具的CodeBERT重新实现得分远高于恒定基线,但项目分开后并不比基线好。高分依赖于标注捷径和评估协议,而非测试代码。在FlakeBench上,模型能识别几乎所有不稳定测试。第二个数据集从CI日志中挖掘,包含86个在同一提交上通过和失败的不稳定端到端测试。测试代码和CI日志能找出42%的原因,其余58%需进一步执行证据。我们重新构建不稳定预测,围绕观察到的失败是否不稳定以及给定执行环境下测试失败的可能性展开。我们的数据集和CI挖掘方法支持这一方向。
英文摘要
Flaky tests pass and fail on the same code version, weakening the signal of test results and disrupting continuous integration (CI) pipelines. Code-based flakiness detectors report strong benchmark results, yet their use in practice remains limited. We argue that the field is studying the wrong problem: Flakiness is not a static property of test code, which often lacks the information needed to decide whether a test is flaky. Analyzing three code-based detectors operating on test code, we found that widely used benchmarks contain shortcuts that inflate reported F1 scores and that evaluation protocols overstate generalizability. To control for these shortcuts, we curated two datasets. The first, C-IDoFT (54,468 unit tests from 57 GitHub projects), keeps a developer-confirmed subset of IDoFT's flaky tests and rebuilds only the non-flaky class from repeated executions instead of fixed versions of flaky tests. C-IDoFT is a controlled counterfactual, not a benchmark for reuse. Our CodeBERT reimplementations of two published detectors scored far above its constant baselines under the published cross-validation protocol but no better than them once projects were separated. The high scores rested on the labeling shortcut and the evaluation protocol, not on the test code. On FlakeBench, a benchmark restricted to flakiness types typically recognizable from test code, and the same project-disjoint protocol, the models identified nearly all flaky tests. The second dataset, mined from CI logs, contains 86 flaky end-to-end tests that passed and failed on the same commit. The test code and CI log yielded a cause for 42% of them; the other 58% required further execution evidence. Rather than abandoning flakiness prediction, we reframe it around whether an observed failure is flaky and how likely a test is to fail given its execution environment. Our datasets and CI-mining method support this direction.
Comments15 pages, 7 figures, under review