收缩评估(经验报告)
Evaluating Shrinking (Experience Report)
浏览论文内容
中文总结 AI 辅助
本文通过对比QuickCheck、Hedgehog、Falsify三种Haskell框架在四个ETNA工作负载下的收缩效果,发现QuickCheck结构收缩更快且反例质量有竞争力,集成收缩无绝对优势,为收缩算法设计提供参考。
中文摘要 AI 辅助
基于属性的测试框架依赖收缩将嘈杂的随机故障转化为开发者可调试的反例。尽管漏洞发现性能被常规测量,但收缩本身很少被定量评估。本文呈现一份关于评估三种Haskell框架(QuickCheck、Hedgehog和Falsify)收缩的经验报告。该比较涵盖四个ETNA工作负载及多个生成器族,包括基于类型、基于API和构造正确的生成器。我们同时测量有效性(使用与穷举搜索找到的真实最小值的树编辑距离)和成本(使用收缩时间及每单位收缩进度的时间)。在这些工作负载中,QuickCheck的结构收缩通常更快,且在最终反例质量上仍具竞争力;集成收缩本身并不保证性能或有效性优势。我们讨论这些结果对收缩算法未来评估与设计的启示。
英文摘要
Property-based testing frameworks rely on shrinking to turn noisy random failures into counterexamples that developers can debug. Although bug-finding performance is routinely measured, shrinking itself is rarely evaluated quantitatively. We present an experience report on evaluating shrinking across three Haskell frameworks: QuickCheck, Hedgehog, and Falsify. The comparison spans four ETNA workloads and several generator families, including type-based, API-based, and correct-by-construction generators. We measure both effectiveness, using tree edit distance to a ground-truth minimum found by exhaustive search, and cost, using shrink time and time per unit of shrinking progress. Across these workloads, QuickCheck's structural shrinking is usually faster and remains competitive on final counterexample quality; integrated shrinking does not by itself guarantee a performance or effectiveness advantage. We discuss what these results imply for future evaluations and designs of shrinking algorithms.