arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.22597stat.MLcs.LG

面向含稀疏模型的稀有事件数据的尺度不变最优采样

Scale-invariant Optimal Sampling for Rare-events Data with Sparse Models

Jing Wang, HaiYing Wang, Qiang Zhang, Hao Helen Zhang

首次发表
浏览论文内容

中文总结 AI 辅助

针对含稀有事件的大规模数据,提出稀疏模型下的尺度不变最优子采样方法,结合 adaptive lasso、IPW 及 MSCL 提升估计效率,通过模拟与真实数据验证性能。

中文摘要 AI 辅助

subsampling(子采样)可有效解决含稀有事件的大规模数据的计算挑战,但过度激进的子采样可能会对估计效率产生不利影响,最优子采样对于缓解信息损失至关重要。不过,现有的最优子采样概率依赖于数据尺度,某些尺度变换可能会导致子采样效率低下。当存在非活跃特征时,该问题更为显著,因为不恰当的尺度变换会任意放大非活跃特征对子采样概率的影响。我们针对这一挑战,在通常假设存在非活跃特征的稀疏模型背景下,提出一种尺度不变最优子采样函数。我们不专注于估计模型参数,而是定义最优子采样函数以最小化预测误差,使用 adaptive lasso(自适应 lasso)来概述估计过程并研究其理论保证。我们首先为稀有事件数据引入 adaptive lasso 估计量,并建立其 oracle(神谕)性质,从而验证子采样的适用性;随后推导了一种尺度不变最优子采样函数,以最小化逆概率加权(IPW)自适应 lasso 的预测误差;最后,我们提出一种基于最大采样条件似然(MSCL)的估计量,以进一步提高估计效率。我们使用模拟数据和真实数据集开展数值实验,以证明所提方法的性能。

英文摘要

Subsampling is effective in tackling computational challenges for massive data with rare events. Overly aggressive subsampling may adversely affect estimation efficiency, and optimal subsampling is essential to mitigate the information loss. However, existing optimal subsampling probabilities depend on data scales, and some scaling transformations may result in inefficient subsamples. This problem is more significant when there are inactive features, because their influence on the subsampling probabilities can be arbitrarily magnified by inappropriate scaling transformations. We tackle this challenge and introduce a scale-invariant optimal subsampling function in the context of sparse models, where inactive features are commonly assumed. Instead of focusing on estimating model parameters, we define an optimal subsampling function to minimize the prediction error, using adaptive lasso to outline the estimation procedure and study its theoretical guarantee. We first introduce the adaptive lasso estimator for rare-events data and establish its oracle properties, thereby validating the use of subsampling. Then we derive a scale-invariant optimal subsampling function that minimizes the prediction error of the inverse probability weighted (IPW) adaptive lasso. Finally, we present an estimator based on the maximum sampled conditional likelihood (MSCL) to further improve the estimation efficiency. We conduct numerical experiments using both simulated and real-world data sets to demonstrate the performance of the proposed methods.

发表机构

  • University of Connecticut(康涅狄格大学)
  • Wills Eye Hospital Thomas Jefferson University(威尔斯眼科医院托马斯杰斐逊大学)
  • University of Arizona(亚利桑那大学)

机构由 AI 辅助整理,请以论文原文为准。

↑