arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

高维M估计中的影响诊断:精确渐近性

Influence Diagnostics in High-dimensional M-estimation: Precise Asymptotics

Hugo Cui

arXiv 2607.09250首次发表:更新:

发表机构

Université Paris-Saclay; CNRS; Laboratoire de mathématiques d’Orsay(巴黎萨克雷大学; 法国国家科学研究中心; 奥赛数学实验室)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究高维M估计中训练点影响,在高维极限\(n\asymp d\)下,刻画高斯设计凸M估计训练集影响分布的极限测度,证明有影响样本靠近决策边界,联系主动学习数据选择启发式方法。

AI 中文摘要

给定训练点对统计模型的影响通常通过留一法影响来衡量,它量化了将其从训练集中移除对模型准确性的影响。在低维、大样本极限\(n\to \infty, d=O(1)\)下,留一法影响的统计特性已被充分理解,但在高维情况下,由于给定样本的影响对所有其他训练样本产生了非平凡的依赖关系,情况变得更加复杂。对于高斯设计下的凸M估计,在高维极限\(n\asymp d\)时,我们表明训练集上影响的分布收敛到一个极限测度,我们对其进行了精确刻画。基于这些结果,我们证明有影响的样本往往靠近决策边界,从而与主动学习中的标准数据选择启发式方法相关。

英文摘要

The impact of a given training point on a statistical model can be measured through its leave-one-out influence on the model parameters, which quantifies how its removal from the training set affects the learned weights. For convex M-estimation under Gaussian design, in the high-dimensional limit $n\asymp d$, we show that the empirical distribution of influences across training points concentrates around a deterministic measure which we sharply characterize. This characterization suggests that influential samples tend to lie on average close to the decision boundary, making contact with a standard data selection heuristic in active learning.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑