arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

你采样即所得:评估网络安全测量中的采样策略

You Get What You Sample: Evaluating Sampling Strategies for Web Security Measurements

Xuenan Zhang, Yuqing Yang, Giancarlo Pellegrino

arXiv 2609.11218首次发表:更新:

发表机构

CISPA Helmholtz Center for Information Security(CISPA 亥姆霍兹信息安全中心)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究首次系统评估网络安全测量中的采样策略,发现Top N采样有偏差,而概率采样更稳健,并提出自适应概率采样建议。

AI 中文摘要

网络测量研究依赖诸如Tranco等域名数据集来量化安全问题的普遍性和影响,但由于高级分析技术的成本,穷举分析这些数据集通常不可行,因此需要使用采样。尽管采样被广泛使用,但其在很大程度上仍受惯例指导——最常见的是选择Top $N$域名——而非基于证据,且其对安全发现有效性和普适性的影响很少得到系统性评估。因此,尚不清楚常见采样策略是否引入系统性偏差、扭曲观察到的漏洞率或限制研究间的可比性。在这项工作中,我们进行了据我们所知首次关于采样方法如何影响测量和结论的全面调查。通过全面的文献综述和对500k个Tranco及24.8M个Common Crawl主机的大规模测量,我们对数据集和采样策略进行了比较评估。我们表明,虽然Top $N$采样可能是一种合理策略,但研究人员必须牢记Top $N$不能反映网络的整体分布。相反,基于概率的策略能为普遍性和许多影响目标提供稳定、无偏的估计。混合采样相比纯概率采样没有优势,因为其确定性前缀始终对准确性产生负面影响。基于这些结果,我们为未来研究提供了数据支持的指导,建议采用自适应基于概率的采样策略,即使在目标问题普遍性未知时也保持有效。

英文摘要

Web measurement studies rely on domain datasets such as Tranco to quantify the prevalence and impact of security issues at scale, but exhaustively analyzing these datasets is often infeasible because of the cost of advanced analysis techniques, requiring the use of sampling. Despite its widespread use, sampling remains largely guided by convention---most commonly \emph{Top $N$} domain selection---rather than evidence, and its influence on the validity and generalizability of security findings has received little systematic evaluation. Consequently, it remains unclear whether common sampling strategies introduce systematic bias, distort observed vulnerability rates, or limit comparability across studies. In this work, we undertake, to the best of our knowledge, the first comprehensive investigation into how sampling methodologies affect the measurements and the conclusions. Through a comprehensive literature review and large-scale measurements of 500k Tranco and 24.8M Common Crawl hosts, we perform a comparative evaluation of datasets and sampling strategies. We show that, while Top $N$ sampling may be a rational strategy, the researchers have to bear in mind that Top $N$ does not reflect the overall distribution of the web. Instead, probability-based strategies yield stable, unbiased estimates for prevalence and many impact objectives. Hybrid sampling provides no advantages over pure probability sampling, as its deterministic prefix consistently contributes negatively to accuracy. Building on these results, we provide data-backed guidance for future studies, proposing to use an adaptive probability-based sampling strategy that remains effective even when the prevalence of the target issue is unknown.

DOI:10.1145/3830454.3846754

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑