AI 中文总结
该研究探讨训练日志能否提升随机训练模型比较的精度,采用针对各模型的协变量调整方法,经视觉实验发现早期训练日志的简单调整可降低比较不确定性,但协变量选择的噪声会限制效果。
AI 中文摘要
比较随机训练的模型需要从多次运行中估计性能差异及其不确定性。我们研究这些运行的训练日志是否能让此类比较更精确。由于训练日志协变量是在训练期间产生而非预先测量的,我们采用针对各模型的协变量调整:每个模型仅用自身运行的统计量进行调整,且原始均值差仍为报告的效应。在一项涉及三种架构和三个数据集的视觉研究中,基于早期训练日志的简单调整常能降低模型比较的不确定性。主要局限在于协变量选择:广泛搜索日志池以寻找最具相关性的统计量时,即便事后看存在有用统计量,也往往会引入多于消除的噪声。因此,训练日志对更精确的模型比较似乎有用,但仅当调整避免大的选择噪声时才成立。
英文摘要
Comparing stochastically trained models requires estimating both a performance difference and its uncertainty from repeated runs. We study whether training logs from those same runs can make such comparisons more precise. Because training-log covariates are produced during training rather than measured before it, we use arm-specific covariate adjustment: each model is adjusted only with statistics from its own runs, and the raw mean difference remains the reported effect. In a vision study spanning three architectures and three datasets, simple adjustments based on early training logs often reduce uncertainty in model comparisons. The main limitation is covariate selection. Broadly searching the log pool for the most correlated statistic often adds more noise than it removes, even when useful statistics exist in hindsight. Training logs therefore appear useful for more precise model comparisons, but only when the adjustment avoids large selection noise.
CommentsAccepted at the ICML 2026 Workshop on Hypothesis Testing
Journal refThe ICML 2026 Workshop on Hypothesis Testing