arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

靠偷看取胜:未强制的预算与测试集选择夸大了短预算AutoML比较

Winning by Peeking: Unenforced Budgets and Test-Set Selection Inflate Short-Budget AutoML Comparisons

Guilin Zhang, Kai Zhao

arXiv 2608.07303首次发表:更新:

AI 中文总结

该研究指出短预算AutoML比较存在协议缺陷,导致Orcetra的胜率被夸大,修正后其胜率显著下降,还给出了选择偏差的实测函数与短预算比较检查表。

AI 中文摘要

在短时间预算(数十秒而非数小时)下对AutoML系统进行比较的情况,在工具README和研讨会论文中很常见,但这类比较很容易出错。我们报告了一个案例研究,其中简单的AutoML引擎Orcetra在513个OpenML数据集上似乎击败了FLAML和AutoGluon:在名义60秒预算下赢得57.1%的数据集,在30秒预算下仅对FLAML就赢得78.4%的数据集。这两个优势都来自结果表无法显示的协议缺陷:搜索循环在测试划分上对每个候选者评分并报告最佳结果,使得头条指标是数十个有噪声估计的最大值,而基线在训练数据上选择且仅接触测试集一次;预算在启动候选者前检查,但在候选者运行期间从未强制,因此系统在60秒预算下的中位消耗为120秒,是AutoGluon实际挂钟时间的2.24倍。将选择移至验证划分、外部强制截止时间并将每个框架固定到机器的相等份额后,Orcetra在重运行子集上的胜率从59.4%降至34.3%,且与两个竞争者的任何成对差异都不再显著。在单次搜索中记录两个估计量可解释胜率的下降:选择规则占4.8个百分点,计算量不平等占其余大部分。相同的轨迹给出了选择偏差作为预算的函数,该偏差是实测而非假设的:它随K增长,但仅达到0.27的准确率点,约为边际标准误差论证预测的σ√(2lnK)界限的五分之一,因为在共享测试行上评分的候选者抵消了大部分噪声。我们最后提供了短预算比较的检查表,代码、每个数据集的结果以及生成论文中所有数字和图表的脚本随论文发布。

英文摘要

Comparisons between AutoML systems at short time budgets -- tens of seconds rather than hours -- are common in tool READMEs and workshop papers, and they are easy to get wrong. We report a case study in which a simple AutoML engine, Orcetra, appeared to beat FLAML and AutoGluon on 513 OpenML datasets, winning 57.1% of them at a nominal 60-second budget and 78.4% of datasets against FLAML alone at 30 seconds. Both margins came from protocol defects that a results table cannot show. The search loop scored every candidate on the test split and reported the best, making the headline metric a maximum over dozens of noisy estimates while the baselines selected on training data and touched the test set once; and the budget was checked before launching a candidate but never enforced during one, so the system consumed a median of 120 s against a 60-second budget, 2.24x the wall-clock AutoGluon used. Re-running with selection moved to a validation split, the deadline enforced externally and every framework pinned to an equal share of the machine, Orcetra's win rate on the re-run subset falls from 59.4% to 34.3% and no pairwise difference against either competitor remains significant. Recording both estimands inside a single search lets us attribute the collapse: the selection rule accounts for 4.8 percentage points and unequal compute for most of the rest. The same traces give the selection bias as a function of budget, measured rather than assumed: it grows with $K$ but reaches only 0.27 accuracy points, about five times below the $σ\sqrt{2\ln K}$ bound a marginal-standard-error argument predicts, because candidates scored on shared test rows cancel most of the noise. We close with a checklist for short-budget comparisons. Code, per-dataset results and the scripts that regenerate every number and figure in the paper are released with it.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑