从单一稳健拟合中提炼的稀疏回归
Sparse Regression Distilled from a Single Robust Fit
浏览论文内容
中文总结 AI 辅助
针对稳健拟合过密或不稳的问题,提出惩罚蒸馏方法,沿坐标下降路径拟合SCAD估计器,在保持预测稳定性的同时实现稀疏化,并在超导数据上验证了其效果。
中文摘要 AI 辅助
稳健线性拟合能够抵抗响应污染,但可能过于密集或不稳定,难以用于有效的全局解释。我们提出惩罚蒸馏方法,该方法在一条受保护的坐标下降路径上,将平滑裁剪绝对偏差(SCAD)估计器拟合到稳健初始估计器的经验拟合表面,并分别评估候选状态在保真度、简约性、扰动稳定性和留出预测方面的表现。新结果将状态关联到算法实际计算的内容。在固定无污染设计的条件下,确定性界将响应替换有界性从初始拟合传递到每个保留的路径状态。转向固定维度,我们通过经验格拉姆投影和影响函数刻画了神谕支撑分支,给出了协方差加权最小二乘近似等价的条件,并建立了一个路径条件广义信息准则。相比之下,在较大的维度样本比下,全坐标稳健拟合会无预警地崩溃,而筛选则恢复了该构造。在确定筛选框架下,稳健性界以及支撑和选择保证传递到筛选后的拟合。模拟实验在维度样本比(p 高达 240)和信号密度上区分了稳健性传递与支撑恢复、效率和计算,这隔离了当筛选过度选择后稀疏阶段所增加的内容。在一项重复分组的超导研究中,蒸馏估计器在预设的训练响应偏移下保持预测稳定性,但保留了 81 个斜率中的 66.8 至 68.8 个。更强的稀疏化仅在可见的保真度和预测代价下将模型减少到 12.6 至 14.0 个斜率。因此,蒸馏在这些数据上保持了预测稳定性,而没有证实一个紧凑的坐标级解释。
英文摘要
Robust linear fits can resist response contamination yet remain too dense or unstable for useful global explanations. We propose penalized distillation, which fits a smoothly clipped absolute deviation (SCAD) estimator to a robust initial estimator's empirical fitted surface along a safeguarded coordinate-descent path and evaluates candidate states separately for fidelity, parsimony, perturbation stability, and held-out prediction. The new results attach to the states the algorithm actually computes. Conditional on a fixed uncontaminated design, deterministic bounds transfer response-replacement boundedness from the initial fit to every retained path state. Turning to fixed dimension, we characterize the oracle-support branch by its empirical-Gram projection and influence function, give conditions for covariance-weighted least-squares approximation equivalence, and establish a path-conditional generalized information criterion. By contrast, at large dimension-to-sample ratios the full-coordinate robust fit collapses without warning, and screening restores the construction. Under a sure-screening framework, the robustness bound and the support and selection guarantees transfer to the screened fit. Simulations separate robustness transfer from support recovery, efficiency, and computation across the dimension-to-sample ratio, with p up to 240, and the signal density, which isolates what the sparse stage adds once the screen over-selects. In a duplicate-grouped superconductivity study, the distilled estimator remains predictively stable under prespecified training-response shifts but retains 66.8--68.8 of 81 slopes. Stronger sparsification reduces the model to 12.6--14.0 slopes only at visible fidelity and prediction cost. Distillation therefore preserves predictive stability on these data without substantiating a compact coordinate-level explanation.
发表机构
- Kangwon National University(江原国立大学)
机构由 AI 辅助整理,请以论文原文为准。