arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.31257eess.SPstat.ME

可扩展的预测变量依赖下的变量选择:自适应虚拟哑变量方法

Scalable Variable Selection under Predictor Dependence with Adaptive Virtual Dummies

Taulant Koka, Michael Muma

首次发表
浏览论文内容

中文总结 AI 辅助

针对高维变量选择中预测变量依赖导致的FDR失控与计算冗余问题,提出基于自适应虚拟哑变量和共享惰性Gram缓存的方法,实现经验FDR控制并获超六倍加速,支持十万预测变量。

中文摘要 AI 辅助

可靠的高维变量选择需要可扩展且能控制错误率的方法。终止随机实验(T-Rex)选择器通过聚合提前终止的前向选择路径来估计错误发现率(FDR),在这些路径中,预测变量与合成哑变量竞争。我们解决了剩余的两个挑战:i) 预测变量依赖可能使哑变量与预测变量的竞争产生偏差;ii) 计算资源浪费在跨实验重复计算共享项上。基于内存高效的虚拟哑变量(这些哑变量通过其精确条件分布顺序采样投影),我们估计非活跃预测变量的条件高斯分布,并将采样映射到剩余的球面半径上。一个共享的惰性Gram缓存使得响应乘积和每个请求的Gram列在跨实验中仅计算一次。在三种协方差结构下的模拟表明,均匀球面哑变量可能超过目标FDR,而所提出的方法经验性地控制了FDR。缓存带来了超过六倍的加速,并且该方法在10万个预测变量的情况下仍然可行,而竞争性的FDR控制方法在此时已变得计算上不可行。

英文摘要

Reliable high-dimensional variable selection requires scalable error-controlling methods. The Terminating-Random Experiments (T-Rex) selector estimates the false discovery rate (FDR) by aggregating early-terminated forward-selection paths in which predictors compete with synthetic dummies. We address two remaining challenges: i) predictor dependence can bias dummy-predictor competition; ii) computation is wasted on recomputing terms shared across experiments. Building on memory-efficient virtual dummies, which sequentially sample projections from their exact conditional law, we estimate the conditional Gaussian law of inactive predictors and map draws onto the remaining sphere radius. A shared lazy Gram cache computes the response product and each requested Gram column once across experiments. Simulations across three covariance structures show that uniform spherical dummies can exceed the target FDR, whereas the proposed method empirically controls FDR. Caching yields more than a sixfold speedup, and the method remains feasible with 100000 predictors, where competing FDR-controlling methods become computationally impractical.

↑