AI 中文总结
研究针对非独立同分布数据集的马尔可夫链变化点检测难题,提出结合马尔可夫链拉德马赫复杂度与惩罚自适应聚类的非参数算法,可准确获取变化点,证明速率紧密性并讨论计算考量。
AI 中文摘要
离线变化点检测旨在在给定数据序列中检测分布变化的时间点,常用于信号处理、语音处理、气候学等领域。然而,对于非独立同分布数据集的严格非参数变化点检测技术仍难以捉摸。本文提出一种非参数聚类算法,通过结合马尔可夫链的拉德马赫复杂度和惩罚自适应聚类,从长度为\(n\)的给定马尔可夫数据集中准确获取变化点。首先利用再生马尔可夫链的拉德马赫复杂度进展推导出马尔可夫链经验分布的杜瓦雷茨基-基弗-沃尔福威茨(DKW)型不等式,进而证明自适应聚类算法能恢复马尔可夫序列的正确变化点,并表明速率紧密性。最后讨论了该问题的计算考量。
英文摘要
Offline change point detection tries to detect time points of distribution change in a given data sequence; and is now routinely used in signal processing, speech processing, climatology etc. Despite this broad applicability across economics, computer science, and planetary sciences, rigorous, nonparametric techniques for change point detection with non-independent and identically distributed (i.i.d.) datasets has remained elusive. This paper establishes such guarantees by proposing a non-parametric clustering algorithm which can accurately obtain the change points from a given Markovian dataset of length $n$. It does so by bridging together two different components of mathematical statistics; Rademacher complexities of Markov chains, and adaptive clustering via penalisation. Our first result uses recent advances in Rademacher complexities of regenerating Markov chains to derive a Dvoretzky Kiefer Wolfowitz (DKW) type inequality for the empirical distribution of the Markov chain. We then use this to show that an adaptive clustering algorithm recovers the correct change points for a Markovian sequence. We establish the tightness of our rates by showing that they essentially coincide with the best known rates for i.i.d. data. We end the paper by discussing the computational considerations of the problem.
CommentsAccepted in AISTATS, 2026