arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

残差网络(ResNets)中作为测试准确率训练时间预测因子的相变频率

Phase Transition Frequency as a Training Time Predictor of Test Accuracy in ResNets

Arunan J

arXiv 2609.05194首次发表:更新:

AI 中文总结

该研究以残差网络微调的类别可分性相变次数为测试准确率预测因子,在多基准实验中验证其分布内预测能力,为低成本训练质量探测提供新指标。

AI 中文摘要

本文实证研究了在残差网络(ResNet)微调过程中观测到的离散类别可分性跳变次数,将其作为最终测试准确率的预测因子。在四个基准数据集(CIFAR-10、CIFAR-100、TinyImageNet、CIFAR-10-C)和三种架构(ResNet-18、ResNet-50、ResNet-101)上开展了75项实验,每个配置设置5至10个随机种子,在标准独立同分布(i.i.d.)分类基准上得到了强数据集内负相关关系:CIFAR-10上的相关系数r=-0.84(p<10^-8,样本量n=30),CIFAR-100上的相关系数r=-0.87(p<10^-5,样本量n=15)。在分布偏移压力下,该关系有所减弱:TinyImageNet的相关系数r=-0.45,CIFAR-10-C损坏基准的相关系数r=-0.19。两项额外分析验证了该实证结论:将架构深度作为线性协变量进行偏相关分析显示,在CIFAR-100上,相变次数仍具有统计显著的预测能力(偏相关系数r_partial=-0.69,p=0.007);在更严格的类别条件下,样本量n=15时未建立对应的显著结果。与6种其他训练曲线信号的对比表明,相变次数在CIFAR-100上达到了评估信号中最强的相关性,在CIFAR-10上也是最强的信号之一,但在两个受压力基准上被其他信号超越。该对比仅局限于训练曲线层面的信号;与有效秩、海森锐度、费舍尔信息、间隔和神经崩溃等当前文献中最强竞争指标的对比不在本研究范围内,仍有待探索。本文将该观测结果作为一类候选探测指标中的分布内训练质量探测指标,并提供了一种可与标准训练循环同步记录的低成本检测流程。

英文摘要

The number of discrete class-separability jumps observed during ResNet finetuning is examined empirically as a predictor of final test accuracy. Across 75 experiments spanning four benchmarks (CIFAR-10, CIFAR-100, TinyImageNet, and CIFAR-10-C) and three architectures (ResNet-18, ResNet-50, and ResNet-101), with five to ten seeds per configuration, a strong within-dataset negative correlation is obtained on standard i.i.d. classification benchmarks: \(r = -0.84\) on CIFAR-10 (\(p < 10^{-8}\), \(n = 30\)) and \(r = -0.87\) on CIFAR-100 (\(p < 10^{-5}\), \(n = 15\)). Under distributional stress, the relationship attenuates: TinyImageNet yields \(r = -0.45\), and the CIFAR-10-C corruption benchmark yields \(r = -0.19\). Two additional analyses discipline the empirical claim. A partial correlation controlling for architecture depth, treated as a linear covariate, shows that on CIFAR-100 the transition count retains statistically significant predictive power (\(r_{\mathrm{partial}} = -0.69\), \(p = 0.007\)); the corresponding result under the stricter categorical conditioning is not established at \(n = 15\). A comparison against six alternative training-curve signals shows that transition count achieved the strongest correlation among the evaluated signals on CIFAR-100 and one of the strongest on CIFAR-10, but is dominated by other signals on the two stressed benchmarks. The comparison is restricted to training-curve-level signals; comparisons against effective rank, Hessian sharpness, Fisher information, margin, and neural-collapse measures, which are the strongest competitors in the current literature, are not part of the present study and remain open. The observation is presented as an in-distribution training-quality probe among a family of candidate probes, and an inexpensive detection procedure suitable for logging alongside a standard training loop is provided.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑