arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.18695stat.ME

gridcp:Python 中的快速在线变点检测

gridcp: Fast Online Changepoint Detection in Python

Per August Jarval Moen, Sebastian Grau Nielsen, Espen Bjørge Urheim, Martin Tveten, Ingrid Kristine Glad

首次发表
浏览论文内容

中文总结 AI 辅助

该研究提出开源 Python 包 gridcp,基于网格方法将离线变点检验转为高效在线检测器,含 Numba 加速内置检验,支持自定义检验与阈值校准,经模拟和真实数据验证其高效准确。

中文摘要 AI 辅助

在线变点检测是实时检测数据流中分布变化的问题。大量方法适用于离线(固定大小)场景,但将这些方法应用于在线场景很快会变得不可行,因为每个观测值的计算成本和内存消耗通常至少随样本量线性增长。最近提出的一种基于网格的方法(Moen,2026)通过在稀疏几何网格的分割点上评估离线检验统计量来克服这一问题,网格点在过去越远间隔越大。对于广泛的检验统计量类别,该方法使更新时间和内存消耗随数据流长度呈对数增长,同时提供检测延迟的有限样本保证。基于此方法,我们提出了 gridcp,一个开源 Python 包,它通过单一统一接口将离线变点检验转化为高效的在线检测器。用户可从九个 Numba 加速的内置检验中选择,涵盖均值、方差、协方差、回归系数的变化,以及非参数检验和指数族模型的广义似然比检验。用户也可提供自己的检验,该包以相同方式处理。对于任何检验,gridcp 提供蒙特卡洛例程,将检测阈值校准至目标虚警概率或平均运行长度,包括无参数零模型时的数据驱动变体。通过模拟和三个真实数据案例研究,我们表明校准准确,运行时间随流长度和维度的扩展良好,从校准到部署的完整流程在真实长数据流上高效运行,且仅具有较短的检测延迟。

英文摘要

Online changepoint detection is the problem of detecting distributional changes in a data stream in real-time. A large body of methodology exists for the offline (fixed-size) setting, but applying these methods online quickly becomes infeasible since the per-observation computational cost and memory consumption typically grow at least linearly with the sample size. A recently proposed grid-based methodology (Moen, 2026) overcomes this by evaluating an offline test statistic over a sparse geometric grid of split points, with grid points spaced increasingly far apart further in the past. For a wide class of test statistics, this approach keeps update time and memory consumption growing logarithmic in the length of the data stream, while admitting finite-sample guarantees on the detection delay. Building on this methodology, we present gridcp, an open-source Python package that turns offline changepoint tests into efficient online detectors through a single, uniform interface. Users can choose from nine Numba-accelerated built-in tests, spanning changes in the mean, variance, covariance, and regression coefficients, as well as nonparametric tests and generalized likelihood-ratio tests for exponential-family models. Users can also supply their own test, which the package handles identically. For any test, gridcp provides Monte Carlo routines that calibrate the detection threshold to a target false alarm probability or average run length, including data-driven variants when no parametric null model is available. Through simulations and three real-data case studies, we show that calibration is accurate, that runtime scales favorably with both stream length and dimension, and that the complete pipeline, from calibration to deployment, runs efficiently on long real-world streams with only a short detection delay.

↑