arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.04667stat.MLcs.LG

基于选择性推理的有理可表达算法自动统计检验及其在特征选择中的应用

Automatic Statistical Test for Rationally Expressible Algorithms by Selective Inference, with Applications to Feature Selection

Teruyuki Katsuoka, Tomohiro Shiraishi, Shuichi Nishino, Ichiro Takeuchi

首次发表
浏览论文内容

中文总结 AI 辅助

研究针对选择性推理需手动推导选择事件的局限,提出AutoSI框架,自动构建选择事件,支持有理函数类算法,可处理现有方法无法解决的lasso方法,实验显示其p值控制I类错误且效能高。

中文摘要 AI 辅助

选择性推理(SI)可为通过算法从数据中选出的假设提供统计有效的p值,校正了因同一数据既用于选择假设又用于检验假设而产生的偏差。然而,为新算法开发SI程序需要专家推导并实现选择事件,即假设被选中的条件。由于需为每种新算法重复这种专门工作,精确SI迄今仅适用于有限类别的算法。我们提出AutoSI框架,从两方面消除该障碍:其一,AutoSI可根据算法的单个操作自动构建选择事件,用户只需将算法编写为普通类NumPy代码,无需手动推导任何内容;其二,AutoSI拓宽了SI可处理的选择事件类别——现有精确方法仅限于数据的线性或二次不等式表征的选择事件,而AutoSI涵盖所有可表示为数据有理函数(多项式之比)的算法。我们证明AutoSI计算的p值在有限样本中完全有效。我们在三种特征选择方法上验证AutoSI,每种方法仅需几十行代码。其中一种方法——通过交叉验证R²选择调参的lasso,无法在现有精确SI框架中处理,AutoSI使其成为可能。在合成数据集和真实数据集上的实验表明,所得p值可将I类错误率(即假阳性率)控制在名义水平,同时保持高检验效能。

英文摘要

Selective inference (SI) provides statistically valid $p$-values for hypotheses selected by applying an algorithm to the data, correcting for the bias that arises when the same data are used both to select and to test a hypothesis. Developing an SI procedure for a new algorithm, however, has required an expert to derive, and then implement, the selection event, i.e., the conditions under which the hypothesis is selected. Repeating this specialized effort for every new algorithm is why exact SI has so far been available for only a narrow class. We propose AutoSI, a framework that removes this barrier in two ways. First, AutoSI constructs the selection event automatically from the algorithm's individual operations, so the user only writes the algorithm as ordinary NumPy-like code and derives nothing by hand. Second, AutoSI broadens the class of selection events SI can handle: existing exact methods are limited to selection events characterized by linear or quadratic inequalities in the data, whereas AutoSI covers any algorithm expressible through rational functions of the data (ratios of polynomials). We prove that the $p$-values computed by AutoSI are exactly valid in finite samples. We demonstrate AutoSI on three feature-selection methods, each written in a few dozen lines of code. One of these methods, the lasso with its tuning parameter selected by cross-validated $R^2$, cannot be handled within existing exact SI frameworks and is made possible by AutoSI. Experiments on synthetic and real datasets show that the resulting $p$-values control the type I error rate (i.e., the false positive rate) at the nominal level while retaining high power.

发表机构

  • Nagoya University(名古屋大学)
  • RIKEN(理化学研究所)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑