发表机构
Trendyol Group; Istanbul Technical University(Trendyol集团; 伊斯坦布尔理工大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文评估五种多列双样本漂移检测器在大规模电商数据上的表现,证明分布式最大均值差异方法在Spark上稳健扩展,并揭示逐维KS检验因统计量饱和而失效的局限。
AI 中文摘要
概念漂移威胁着生产环境中的机器学习,然而多变量双样本漂移检测器在大规模场景下的实证行为仍未得到充分刻画。现有基准很少涉及工业运营数据集典型的上亿行数据和高基数特征。我们在三个互补环境中评估了五种多列双样本检验(边际方法、基于投影的方法和核嵌入方法):哈佛数据仓库、一个经过验证的Failing Loudly复现(平均绝对误差在0.030至0.053之间),以及一个在包含1.375亿行的Trendyol集合排序特征表上的新型合成注入基准。我们测试了四种漂移类型,跨越两种严重程度-范围机制,证明了在Apache Spark上使用随机傅里叶特征的分布式最大均值差异方法能够稳健扩展。在强机制下,对四种漂移类型取平均,并在校准阈值下,该方法与预期漂移幅度达到皮尔逊相关系数r=0.940,真阳性率为80.4%,假阳性率为3.2%。相反,逐维Kolmogorov-Smirnov检验因ID类列在非对称采样下导致的统计量饱和而失败,这为大规模采样设计确立了一个关键约束。在弱配置下(实现翻转比例最多为0.57%),检测器难以可靠区分,凸显了未来需要强度网格功效分析,以区分基本灵敏度界限与可扩展的阈值偏移。
英文摘要
Concept drift threatens production machine learning, yet the empirical behavior of multivariate two-sample drift detectors at scale remains under-characterized. Existing benchmarks rarely address the hundreds of millions of rows and high-cardinality features typical of industrial-operational datasets. We evaluate five multi-column two-sample tests (marginal, projection-based, and kernel embedding methods) across three complementary environments: the Harvard Dataverse, a validated Failing Loudly reproduction (mean absolute error between 0.030 and 0.053), and a novel synthetic-injection benchmark on the 137.5-million-row Trendyol collection-ranking feature table. Testing four drift types across two severity-scope regimes, we demonstrate that distributed Maximum Mean Discrepancy with Random Fourier Features on Apache Spark scales robustly. Averaged over the four drift types in the strong regime and under a calibrated threshold, it achieves a Pearson correlation of r = 0.940 with expected drift magnitude, an 80.4% true positive rate, and a 3.2% false positive rate. Conversely, the per-dimension Kolmogorov-Smirnov test failed due to statistic saturation from ID-like columns under asymmetric sampling, establishing a critical constraint for large-scale sampling design. At weak configurations (realized-flip fractions of at most 0.57%), detectors struggled to reliably discriminate, highlighting the need for future intensity-grid power analyses to distinguish fundamental sensitivity bounds from scalable threshold shifts.