通过渐进采样进行大规模数据质量分析:以数据为中心的人工智能管道的基准测试
Data Quality Profiling at Scale with Progressive Sampling: A Benchmark for Data-Centric AI Pipelines
浏览论文内容
中文总结 AI 辅助
研究以数据为中心的人工智能管道中数据质量分析问题,通过对九种采样策略在多类数据集上基准测试发现,盲目代表性采样器效果最佳,代理引导方法欠佳,根本原因是代理不匹配,随机均匀或聚类采样可满足大规模生产级质量分析。
中文摘要 AI 辅助
数据质量分析对于以数据为中心的人工智能管道至关重要,包括计算缺失值率、重复率、异常值密度和函数依赖违规情况等。然而,对数百万行数据进行详尽扫描对于近实时监控来说速度过慢。渐进采样是标准替代方法,但哪种策略在大规模下能最佳保持分析保真度仍是开放问题。我们在三个真实世界数据集、一个物联网传感器流、两个超大型真实数据集以及合成数据上对九种采样策略进行基准测试。结果表明,盲目代表性采样器在5%预算下表现最佳,代理引导方法存在问题,根本原因是代理不匹配,随机均匀采样或聚类采样足以满足大规模生产级质量分析。
英文摘要
Data quality profiling -- computing missing-value rates, duplicate fractions, outlier densities, and functional-dependency violations -- is foundational for data-centric AI pipelines, yet exhaustive scans over millions of rows are prohibitively slow for near-real-time monitoring. Progressive sampling is the standard alternative; the open question is which strategy best preserves profile fidelity at scale. We benchmark nine sampling strategies -- blind (random uniform, geometric, Yamane, cluster) and proxy-guided (Metropolis-Hastings, DAG, stratified by column type or quality score, importance-weighted) -- on three real-world datasets (NYC 311, NYPD arrests, UCI Adult; up to 500K rows), an IoT sensor stream (2.3M rows), two ultra-large real datasets including Ultra-Marathon Running (up to 7.4M rows), and synthetic data scaled to 5x10^6 rows. Contrary to the assumption sharpens estimates, blind representative samplers dominate uniformly. At a 5% budget, random uniform achieves 0.49% mean relative error on NYC 311; DAG-guided MCMC yields 19.5% (approx. 40x worse), and across all real datasets DAG is 11-49x worse (Wilcoxon W=0, p=0.002, n=9 pairs). Cluster sampling matches random uniform (MRE 0.110 vs. 0.111); proxy-guided methods share DAG's failure mode (MRE 0.20-0.35). At scale, random uniform is near-linear (O(N^{0.964})) while DAG is super-linear (O(N^{1.272})), running 28--47x slower on ultra-large data with 6x worse accuracy. The root cause is an IQR proxy mismatch: proxy-guided samplers over-pursue numeric outliers, while quality defects concentrate in categorical columns invisible to the proxy. The actionable finding: representativeness, not domain knowledge, determines sampler quality -- schema-free random uniform or cluster sampling suffices for production-grade quality profiling at scale.
发表机构
- IRD(法国发展研究院)
- ESPACE-DEV(ESPACE-DEV机构)
机构由 AI 辅助整理,请以论文原文为准。