发表机构
Delft University of Technology(代尔夫特理工大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究实证检验了在测试集上调超参数的危害,发现其性能膨胀虽真实但常较小,且模型排名基本不变,呼吁更细致看待并公开报告测试调优。
AI 中文摘要
机器学习教科书经常告诫“不要在测试集上调整超参数”。违反这一原则被视为一种严重的错误,会产生误导性的乐观结果,破坏基准测试的完整性,甚至可能被解读为科学欺诈。然而,有证据表明,在测试集上进行超参数调优在实践中确实存在,这使得理解其实际后果变得更加重要。那么,它到底有多糟糕呢?在这项工作中,我们质疑这一教条,并对其进行实证检验。我们系统地研究了在MNIST-1D、CIFAR-10以及GLUE基准测试中的三个任务上,在测试集上调整超参数所导致的性能膨胀程度。我们的实验表明,虽然这种影响是真实且显著的,但相对于其他噪声来源,它通常较小。在许多情况下,我们发现,在测试集上调优与在验证集上调优所得到的模型完全相同。最重要的是,我们发现,在测试集上调优后,模型的排名基本保持不变,因此,一致的测试集调优可能不会使基准测试或模型选择失效。我们的结果呼吁对在测试集上调整超参数采取更细致的看法,鼓励研究人员公开报告测试调优。
英文摘要
"Don't tune hyperparameters on the test set" is often stated in machine learning textbooks. Violating it is considered a cardinal sin that produces misleadingly optimistic results, corrupts benchmark integrity, and thus can even be interpreted as scientific fraud. Yet evidence suggests that test set hyperparameter tuning does occur in practice, making it all the more important to understand its actual consequences. So how bad is it, really? In this work we question this dogma and put it to an empirical test. We systematically study the magnitude of the performance inflation caused by tuning the hyperparameters on the test set for MNIST-1D, CIFAR-10, and three tasks from the GLUE benchmark. Our experiments show that while the effect is real and significant, it is frequently small relative to other sources of noise. In many cases, we find that tuning on the test set recovers exactly the same model as when tuning on the validation set. Most importantly, we find that the rankings of models remain essentially preserved after tuning on the test set and therefore that consistent test-set tuning may not invalidate benchmarks or model selection. Our results call for a more nuanced view of tuning hyperparameters on the test set, stimulating researchers to openly report test tuning.