AI 中文总结
研究用高斯混合模型检测异常与分布漂移,明确三个实际选择并评估。通过自动选择混合成分数量、负对数似然评分及扩展模型检测漂移,在多基准测试中该模型有竞争力且可解释,还对比了其他方法。
AI 中文摘要
我们重新审视高斯混合模型(GMMs),将其作为一种轻量级、可解释的异常检测工具,特别是用于检测数据流中的分布漂移。我们明确了三个实际选择,并在七个公共基准上进行评估。首先,混合成分的数量由贝叶斯信息准则自动选择,通过k-means初始化,无需预先固定。其次,根据拟合正常数据的GMM下的负对数似然对单个观测进行评分,使用极值理论将阈值设置为目标误报率。第三,相同的可解释模型扩展到分布漂移:每个高斯成分是一个命名的“状态”,与任何状态都不匹配的流窗口部分——其未解释质量——是一个漂移信号,本身就是解释。我们将此与无模型核两样本检验(最大均值差异,MMD)以及两种GMM到GMM的散度(闭式柯西-施瓦茨散度和基于匹配的KL替代)进行基准测试。在从3到64维的七个基准和五次随机分割中,GMM点检测器与隔离森林、局部离群因子、一类支持向量机、ECOD、COPOD和自动编码器具有竞争力,虽然很少比它们更准确,但独特地产生了一个可解释的模型。对于漂移,MMD是最强的纯检测器,但当异常形成新状态时,可解释的未解释质量统计与它匹配(并且诚实地失败,因为MMD在漂移是现有状态的纯重新加权时不会失败)。每个警报都是可解释的:异常位于其最近状态之外3-10个标准差的中位数处,而正常点约为1个标准差,漂移警报报告与任何已知状态都不匹配的窗口部分。所有代码和实验都已发布。
英文摘要
Drift detectors that work tend not to explain themselves, and drift detectors that explain themselves tend to fail in high dimension. We close that gap for Gaussian mixture models (GMMs): each fitted component is a named "regime," and the fraction of a stream window matching no regime -- its unexplained mass -- is a drift signal that is simultaneously its own explanation. We identify why this statistic collapses in high dimension and repair it. Under a correct component a normal point in d dimensions lies about sqrt(d) sigma from the mean, so once d exceeds 9 essentially every point exceeds a fixed 3-sigma radius: window-level ROC-AUC is exactly 0.50 on Satellite (d=36) and Optdigits (d=64). Calibrating the radius to sqrt(chi-squared_d(0.99)) removes the collapse -- AUC 1.00 and 0.89 -- while leaving low dimensions unchanged. Across seven public benchmarks, five seeds, and eight model-free detectors spanning the kernel, classifier, projection, density-difference, transport, likelihood and partition families, the repaired statistic is best or tied-best on five of seven datasets at 10% window contamination (its two losses are Pendigits, where the whole field beats it, and Optdigits), and as contamination becomes sparse the sample-level detectors fade toward chance while it degrades most gracefully: at 2% its mean AUC across the benchmarks is 0.86 against at most 0.73 for any model-free detector (1.00 vs. MMD's 0.72 on KDD-http) -- while alone among them reporting which regime the data left and how far outside it the window lies. We delimit its scope honestly: unexplained mass detects and explains novel-regime drift but is blind by construction to in-support re-weighting of known regimes, where distribution-level tests are required and explain nothing; and the underlying density model's EVT-calibrated false-alarm rates degrade above d of about 36. All code and experiments are released.
Comments26 pages, 5 figures, 9 tables. Code to reproduce all experiments included as ancillary files