一种基于熵的、可校正优化偏差的决定系数
An Entropy-based Coefficient of Determination with Adjustment of Optimization Bias
浏览论文内容
中文总结 AI 辅助
该研究提出基于熵方差的R²指标,校正优化偏差,可用于变量选择,在模拟中使LASSO路径的错误发现率从80%降至6%,且保留信号召回率。
中文摘要 AI 辅助
经典似然比检验和ΔAIC会随样本量扩大而加剧统计显著性危机,常将微不足道的改进标记为高度显著。尽管平均处理效应(ATE)等因果估计量可量化实际效应大小,但它们依赖期望算子,与数据的原始坐标尺度绑定。此外,现有伪R²指标存在不足:基于方差的度量忽略高阶分布变化,且现有公式不具备单调变换不变性。我们通过引入熵方差(EV)作为普通最小二乘中误差方差的严格、与尺度无关的泛化,解决这些局限。我们定义基于总体EV的参数ρ²_V,其将无界交叉熵投影到标准化的[0,1]尺度,并证明基于EV的F_V统计量渐近服从F分布。基于这些分布特性,我们提出两个估计量:经验总体R²_SV和样本外预测R²_SVP。两者均通过对每个观测值的交叉熵取指数得到,并纳入训练乐观度的自由度校正。利用F_V分布,我们为ρ²_V推导了精确的p值和置信区间,无需难以处理的费希尔信息矩阵。模拟研究和帕金森病微生物组应用表明,通过这些EV-R²指标进行变量选择具有优越性。值得注意的是,在模拟中通过数据拆分评估LASSO路径的R²_SVP,将错误发现率从80%降至6%,同时完全保留了信号召回率。
英文摘要
Classical likelihood-ratio tests and $Δ$AIC exacerbate the statistical significance crisis by scaling with sample size, often flagging negligible improvements as highly significant. While causal estimands like the average treatment effect (ATE) quantify practical magnitude, their reliance on the expectation operator ties them to the data's original coordinate scale. Furthermore, existing pseudo-$R^2$ metrics are inadequate: variance-based measures ignore higher-order distributional changes, and current formulations lack invariance to monotone transformations. We resolve these limitations by introducing Entropic Variance (EV) as a rigorous, scale-independent generalization of error variance in ordinary least squares. We define the population EV-based parameter, $ρ^2_V$, which projects unbounded cross-entropy onto a standardized $[0,1]$ scale, and establish that the EV-based $F_\text{V}$ statistic asymptotically follows an $F$-distribution. Building on these distributional properties, we propose two estimators: the empirical population $R^2_{\text{SV}}$ and the out-of-sample predictive $R^2_{\text{SVP}}$. Both are derived by exponentiating per-observation cross-entropy and incorporate a degrees-of-freedom correction for training optimism. Leveraging the $F_\text{V}$-distribution, we derive refined $p$-values and confidence intervals for $ρ^2_V$ without requiring intractable Fisher information matrices. Simulation studies and a Parkinson's disease microbiome application demonstrate the superiority of variable selection via these EV-$R^2$ metrics. Notably, evaluating the $R^2_{\text{SVP}}$ of a LASSO path via data-splitting reduced false discovery rates from 80% to 6% in simulations while fully preserving signal recall.