AI 中文总结
该研究利用深度神经网络早期训练的遥测数据结合超参数,通过梯度提升树等方法,实现对训练最终准确率、相对性能及故障的高准确度预测,可为计算分配提供决策支持。
AI 中文摘要
深度神经网络的大规模超参数搜索会在那些从最初几个轮次就注定无法成功的配置上消耗大量计算资源。本研究探讨:单次训练运行的早期遥测数据——包括每轮损失、训练准确率、梯度信噪比、权重范数增长以及激活饱和快照,结合其采样得到的超参数,能否在不参考其他运行结果的情况下,预测该次运行的最终结果。我们评估三项预测任务:最终测试准确率、同一领域内的相对性能,以及训练动态故障(包括数值发散)。在涵盖6种架构/数据集组合的23788次训练运行中,仅使用前5轮遥测数据的梯度提升树,在永久留出的超参数配置测试集上,最终准确率回归的R²达到0.92-0.99,相对分类的ROC-AUC达到0.983-0.998。仅1轮训练后即可获得有效的预测结果。配对 ablation 实验显示,梯度和权重层级的遥测数据相比仅使用损失和准确率曲线,能提供统计上一致的性能提升,尽管实际增益因领域而异。相似架构间的迁移性较强,而跨数据集迁移的限制主要源于准确率尺度的差异,而非底层关系的丢失。这些结果表明,早期训练遥测数据可作为计算分配的实用决策支持信号,同时需对任何自动化干预进行人工监督。
英文摘要
Large hyperparameter sweeps for deep neural networks spend substantial compute on configurations that are effectively doomed from the first few epochs. We study whether a single training run's own early telemetry - per-epoch loss, training accuracy, gradient signal-to-noise ratio, weight-norm growth, and an activation-saturation snapshot - together with its sampled hyperparameters, can predict that run's eventual outcome without reference to other runs. We evaluate three prediction tasks: final test accuracy, relative performance within a domain, and training-dynamics failure, including numerical divergence. Across 23,788 training runs spanning six architecture/dataset combinations, gradient-boosted trees using only the first five epochs of telemetry achieve R^2 = 0.92-0.99 for final-accuracy regression and ROC-AUC = 0.983-0.998 for relative classification on a permanently held-out set of hyperparameter configurations. Useful prediction is already available after a single epoch. A paired ablation shows that gradient- and weight-level telemetry provides a statistically consistent improvement over loss and accuracy curves alone, although the practical gain varies by domain. Transfer is strong between similar architectures, while cross-dataset transfer is limited mainly by differences in accuracy scale rather than loss of the underlying relationship. These results suggest that early-training telemetry can provide a practical decision-support signal for compute allocation while motivating human oversight for any automated intervention.
Comments21 pages, 6 figures, 7 tables, includes appendices