arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

当漂移检测器发出错误警报时:连续机器学习监测中的误报率

When Drift Detectors cry Wolf: False Alarm Rates in continuous ML Monitoring

Raj Shekhar Singh

arXiv 2607.17336首次发表:更新:

发表机构

Indian Institute of Technology, Roorkee(印度理工学院鲁尔基分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究连续机器学习监测中常用漂移检测器的误报率,通过实证分析PSI、KS等五种检测器,发现不同检测器对批量大小的敏感性不同,应用邦费罗尼校正有稳定性 - 敏感性权衡,为生产ML系统选择和校准检测器提供实用指南。

AI 中文摘要

漂移检测是生产机器学习监测系统的核心组件,用于将传入数据与参考分布进行比较并在变化发生时触发警报。然而,这些检测器在研究中常被评估在合成变化下的检测准确性,而忽视了连续监测下的误报。在生产环境中,模型随时间和众多特征被反复监测,即使小的误报率也会累积成频繁警报,导致警报疲劳。我们实证分析了五种常用漂移检测器(PSI、KS、MMD、LSDD和对抗验证)的误报行为。与现有文献一致,PSI对批量大小敏感,小样本量时频繁误报,但批量超过约200个样本时行为稳定且改善。KS、MMD和LSDD在不同批量大小下有持续波动,在低数据情况下比PSI更可靠。应用邦费罗尼校正降低了误报率,但常以降低真阳性敏感性为代价,强化了漂移检测中稳定性 - 敏感性的权衡。这项工作对连续监测条件下多个漂移检测器的误报行为进行了系统比较,识别了不同检测器系列的权衡,并为生产机器学习系统中选择和校准漂移检测器提供了实用指南。

英文摘要

Drift detection is a core component of production machine learning monitoring systems, where detectors are used to compare incoming data with a reference distribution and trigger alerts when changes occur. However, these detectors are often evaluated in research settings that emphasize detection accuracy under synthetic shifts, while overlooking false alarms under continuous monitoring. In production environments, models are monitored repeatedly over time and across many features, and even small false positive rates can accumulate into frequent alerts, leading to alarm fatigue. We empirically analyze false positive behavior across five commonly used drift detectors: PSI, KS, MMD, LSDD, and adversarial validation. Consistent with existing literature, PSI exhibits strong sensitivity to batch size, producing frequent false alarms at small sample sizes; however, we further observe that its behavior stabilizes and improves substantially once batch sizes exceed approximately 200 samples. In contrast, KS, MMD, and LSDD display persistent fluctuations across batch sizes, while remaining comparatively more reliable than PSI in low-data regimes. Applying a Bonferroni correction reduces false positive rates, but often at the cost of reduced true positive sensitivity, reinforcing the well-known stability - sensitivity trade-off in drift detection. This work provides a systematic comparison of false positive behavior across multiple drift detectors under continuous monitoring conditions. We identify tradeoffs across detector families and provide practical guidelines for selecting and calibrating drift detectors in production ML systems.

CommentsAccepted in ICLR 2026 CAO workshop

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑