arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于分组高斯镜像与排列SHAP的模型不可知FDR控制

Model-Agnostic FDR Control via Group Gaussian Mirror and Permutation SHAP

Jiaan Han, Junxiao Chen, Yanzhe Fu

arXiv 2608.00989首次发表:更新:

发表机构

Columbia Business School; Columbia University; University of Hong Kong(哥伦比亚商学院; 哥伦比亚大学; 香港大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究针对序列与分组模型的FDR控制问题,提出模型不可知的分组特征FDR控制框架,结合分组高斯镜像与排列SHAP,在模拟与真实数据集上实现可靠FDR控制并提升功效。

AI 中文摘要

大多数FDR控制的特征选择方法针对逐坐标假设设计,每个特征有单一权重或重要性得分,该抽象不适用于序列模型和分组模型,其中一个原始特征由子特征块(如滞后项、循环状态或基于注意力的交互)表示。我们为此类场景提出分组特征FDR控制框架:针对分组线性模型,构建带矩阵值扰动的零对称块级镜像统计量;针对神经序列模型,将排列SHAP导数作为模型不可知的块级重要性得分,结合基于核的依赖度量。该框架对网络架构具有模型不可知性,无需指定协变量分布,块大小为1时可退化为高斯镜像或神经高斯镜像。我们证明了低维和高维分组线性模型的FDR控制,以及固定拟合非线性模型下平滑排列SHAP导数的渐近对称性。在模拟和真实世界数据集上的实验表明,该方法在相关分组特征信号下能实现可靠的FDR控制并提升功效。

英文摘要

Most FDR-controlled feature selection methods are designed for coordinate-wise hypotheses, where each feature has a single weight or importance score. This abstraction fails in sequential and grouped models, where one original feature is represented by a block of sub-features, such as lags, recurrent states, or attention-based interactions. We propose a grouped-feature FDR control framework for such settings. For grouped linear models, we construct null-symmetric block-level mirror statistics with matrix-valued perturbations. For neural sequential models, we combine Permutation SHAP derivatives as model-agnostic block-level importance scores with kernel-based dependence measure. The framework is model-agnostic across network architectures, does not require specifying the covariate distribution, and reduces to Gaussian Mirror or Neural Gaussian Mirror when the block size is one. We prove FDR control for low- and high-dimensional grouped linear models and asymptotic symmetry of smoothed Permutation SHAP derivatives under fixed fitted nonlinear models. Experiments on simulated and real-world datasets show reliable FDR control and improved power under correlated grouped-feature signals.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑