AI 中文总结
针对高维多重变点分析,现有方法有局限。本文提出通用U统计框架,通过移动窗口内的两样本U统计量,灵活选核函数。在检验和估计上有创新方法,经实验验证性能优,在基因组数据应用中凸显实用性,还有公开R包。
AI 中文摘要
高维变点分析在现代统计推断中至关重要。现有方法常针对特定参数或特定任务设计,难以推广,且依赖严格分布假设,对重尾数据鲁棒性差。本文提出统一框架用于高维数据中多重变点的检验、估计和推断。利用移动窗口内的两样本U统计量,灵活选择核函数。检验方面,基于L无穷范数统计量和高维乘子自助法,在稀疏假设下达到极小极大最优功效;估计方面,构建初始估计并通过U统计投影细化算法优化,获得极小极大最优定位率,还推导了渐近分布以构建有效置信区间。数值实验表明该方法在多种情况下性能更优,在基因组拷贝数变异数据中的应用凸显其实用性,且有R包可供使用。
英文摘要
High-dimensional change-point analysis is essential in modern statistical inference. However, existing methods are often designed either for specific parameters (e.g., mean or variance) or for particular tasks (e.g., testing or estimation), making them difficult to generalize. Moreover, they typically rely on restrictive distributional assumptions, limiting their robustness to heavy-tailed data. We propose a unified framework for testing, estimating, and inferring multiple change points in high-dimensional data. Our approach leverages a two-sample U-statistic within a moving window, allowing flexible kernel function selection to accommodate structural changes in general parameters such as variance changes or robust statistics. For testing, we develop an L-infinity norm-based statistic with a high-dimensional multiplier bootstrap procedure, achieving minimax-optimal power under sparse alternatives. For estimation, we construct an initial estimator for the change-point number and locations and refine it using the U-statistic Projection Refinement Algorithm (U-PRA), attaining minimax-optimal localization rates. We further derive the asymptotic distribution of refined estimators, enabling valid confidence interval construction. Extensive numerical experiments demonstrate the better performance of our method across various settings, including heavy-tailed distributions. Applications to genomic copy number variation data highlight its practical utility. An R package implementing the proposed method, U-PRA, is publicly available at https://github.com/liubin0145/R-codes-UPRA/.
Comments135 pages