arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

校准的高斯过程变量选择中的导数过程灵敏度

Calibrated Derivative-Process Sensitivity for Gaussian-Process Variable Selection

Jia Cai

arXiv 2609.33549首次发表:更新:

发表机构

George Mason University(乔治梅森大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对高斯过程变量选择中ARD和导数灵敏度缺乏校准规则的问题,提出学生化乘子自助法校准导数过程灵敏度,实现FDR控制,减少虚假输入,成本低。

AI 中文摘要

自动相关性确定(ARD)是高斯过程(GP)回归中变量选择的默认工具,它通过逆长度尺度对输入进行排序——这些尺度衡量函数变化的速度,而非输入对预测的贡献程度——并且没有提供校准的规则来决定保留哪些输入。以预测为中心的替代方案,即导数灵敏度 $\nu_j = \mathbb{E}[(\partial f/\partial x_j)^2]$,可以从拟合的GP中以封闭形式获得,但将其转化为选择规则比看起来更难:在零输入处,估计量是退化的二次型,因此Wald和Bernstein-von Mises截断是反保守的,而自然的残差自助法(residual bootstrap)的尺度是错误的。我们证明,对GP导数过程进行学生化乘子自助法(studentized multiplier bootstrap)可以修复这两个问题,并通过二次型的不变性原理证明其有效性,并在输入间获得渐近的族wise误差和错误发现率控制。在100次重复中,该规则在输入确实为零时控制了FDR,而未校准的导数排名违反目标高达2倍,Bernstein-von Mises截断违反2.2倍;在匹配的FDR下,它不损失功效;在Matérn核和输入相关性高达0.99的情况下依然成立;在具有植入和真实零输入的真实数据上,它减少了5-12倍的虚假输入;其成本为GP拟合的5-18%;并且块平均变体在成本随$n$线性增长的情况下保持有效性。

英文摘要

Automatic relevance determination (ARD), the default tool for variable selection in Gaussian-process (GP) regression, ranks inputs by inverse lengthscales -- which measure how fast a function varies, not how much an input contributes to prediction -- and offers no calibrated rule for deciding which inputs to keep. The prediction-centred alternative, the derivative sensitivity $ν_j = \mathbb{E}[(\partial f/\partial x_j)^2]$, is available in closed form from a fitted GP, but turning it into a selection rule is harder than it looks: at a null input the estimator is a degenerate quadratic form, so Wald and Bernstein-von Mises cutoffs are anti-conservative, and the natural residual bootstrap is mis-scaled. We show that a studentized multiplier bootstrap of the GP derivative process repairs both, prove its validity through an invariance principle for quadratic forms, and obtain asymptotic family-wise and false-discovery-rate control across inputs. Over 100 replications the rule controls FDR wherever inputs are truly null, while uncalibrated derivative rankings breach the target by up to 2x and a Bernstein-von Mises cutoff by 2.2x; at matched FDR it loses no power; it holds under a Matérn kernel and input correlation up to 0.99; on real data with planted and authentic null inputs it admits 5-12x fewer spurious inputs; it costs 5-18% of the GP fit; and a block-averaged variant retains validity at cost linear in $n$.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑