arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.15450math.STcs.LGstat.MLstat.TH

线性回归和逻辑回归中的仅预测蒸馏

Prediction-Only Distillation in Linear and Logistic Regression

  • University of Texas(德克萨斯大学)

机构由 AI 辅助整理,请以论文原文为准。

Hien Dang, Pratik Patil, Alessandro Rinaldo

AI总结:

研究在仅能获取训练好的预测器和未标记协变量时的自蒸馏,通过新方案训练学生模型,推导岭回归最优混合预测风险的确定性等价物,证明其风险更小,还表明最优混合权重可估计,且二元逻辑回归中预测混合能有更好表现。

AI中文摘要:

自蒸馏(SD)通常是在学生模型基于教师模型的原始训练输入进行重新训练时进行研究。然而,在许多实际部署中,标记的训练数据不再可用,只能访问训练好的预测器和新的未标记协变量。我们通过一种新的X预测混合方案在这种仅预测的情况下研究SD:一个纯蒸馏的学生模型在由教师模型伪标记的新协变量上进行训练,最终预测器是教师模型和学生模型预测的仿射组合。对于比例渐近下的岭回归,我们推导了一般各向异性协方差和确定性信号下最优混合预测风险的确定性等价物。我们表明,对于几乎每一对教师和学生正则化水平,这种风险都严格小于教师模型的风险,包括当新协变量分布外甚至其协方差是各向同性时。我们进一步表明,最优混合权重不能仅从未标记数据中识别,但可以在单个训练后步骤中使用一个小的独立标记校准集进行一致估计,而无需额外的模型拟合。最后,对于二元逻辑回归,我们表明预测混合可以优于教师模型和纯蒸馏分类器。

英文摘要:

Self-distillation (SD) is typically studied when the student is retrained on the teacher's original training inputs. In many practical deployments, however, the labeled training data are no longer available, and one has access only to the trained predictor and fresh unlabeled covariates. We study SD in this prediction-only regime through a fresh-X prediction-mixed scheme: a pure-distilled student is trained on fresh covariates pseudo-labeled by the teacher, and the final predictor is an affine combination of the teacher and student predictions. For ridge regression under proportional asymptotics, we derive deterministic equivalents for the optimally mixed prediction risk under general anisotropic covariance and deterministic signal. We show that this risk is strictly smaller than the teacher risk for almost every pair of teacher and student regularization levels, including when the fresh covariates are out-of-distribution and even when their covariance is isotropic. We further show that the optimal mixing weight cannot be identified from unlabeled data alone, but can be consistently estimated in a single post-training step using a small independent labeled calibration set, without additional model fitting. Finally, for binary logistic regression, we show that prediction mixing can outperform both the teacher and the pure-distilled classifier.

补充信息

↑