arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.13922cs.LGstat.ML

基于低秩二阶密度投影的高维非参数变点检测

High-dimensional nonparametric changepoint detection via low-rank degree-two density projection

Guoqing Zhang, Zhaixin Chen

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出基于低秩二阶密度投影的高维非参数变点检测方法,利用低秩矩阵CUSUM等技术实现变点精确检测,在高维数据实验中表现优异。

中文摘要 AI 辅助

当变化前和变化后的密度均未参数化指定时,高维分布变化的检测十分困难。本文提出一种基于表示的方法,该方法保留所有至多二阶的密度信息,同时用矩阵均值估计替代密度估计。对于取值在 $[-1,1]^d$ 的观测值,构造对称特征矩阵 $H_2(X)\in\mathbb{R}^{(d+1)\times(d+1)}$,使得 $M(f)=\mathbb{E}_f H_2(X)$ 是密度二阶正交投影的等距编码。在秩-$r$ 截断后扫描矩阵累积和(matrix CUSUM),利用投影跳变的低秩特性而非单个坐标的稀疏性。所得低秩(LRD)估计量具有帐篷型总体目标函数和非渐近算子范数分析,其主导随机项尺度为 $\sqrt{rd\log(nd)}$。对于多个变点,本文提出带种子的阈值内窄化程序,并通过归纳法证明精确恢复,该归纳法为每个未检测到的变点保留隔离区间。交叉拟合标量细化在一个折上学习变化的低秩方向,在另一个折上定位,达到 $\widetilde{O}_{\mathbb{P}}(\kappa^{-2})$ 的误差;匹配的 Le Cam 下界表明其在对数项范围内是最优的。几何 $\beta$-混合扩展由依赖矩阵伯恩斯坦不等式推导得出。环境维度高达200、含3个变点的 $d=100$ 序列以及128特征的人类活动基准实验表明,该方法计算上仍具实用性,且能准确检测均值 CUSUM 无法察觉的纯依赖关系变化。

英文摘要

Detecting distributional changes in high dimension is difficult when neither the pre-change nor post-change density is parametrically specified. We introduce a representation-based approach that retains all degree-at-most-two density information while replacing density estimation by matrix mean estimation. For observations in $[-1,1]^d$, a symmetric feature matrix $H_2(X)\in\R^{(d+1)\times(d+1)}$ is constructed so that $M(f)=\E_f H_2(X)$ is an isometric encoding of the degree-two orthogonal projection of the density. We scan matrix CUSUMs after rank-$r$ truncation, exploiting the low rank of the projected jump rather than sparsity of individual coordinates. The resulting \LRD{} estimator has a tent-shaped population objective and a nonasymptotic operator-norm analysis whose leading stochastic term scales as $\sqrt{rd\log(nd)}$. For multiple changes, we give a seeded narrowest-over-threshold procedure and prove exact recovery by an induction that preserves an isolating interval for every undetected change. A cross-fitted scalar refinement learns the changing low-rank direction on one fold and localizes on the other, attaining $\widetilde O_{\Pp}(κ^{-2})$ error; a matching Le Cam lower bound shows optimality up to logarithms. A geometrically $β$-mixing extension follows from a dependent matrix Bernstein inequality. Experiments with ambient dimension up to $200$, a three-change $d=100$ sequence, and a $128$-feature human-activity benchmark show that the method remains computationally practical and accurately detects pure dependence changes that are invisible to mean CUSUMs.

发表机构

  • North Carolina State University(北卡罗来纳州立大学)
  • Georgia Institute of Technology(佐治亚理工学院)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑