发表机构
Shandong University(山东大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究通过端到端基准测试和预算审计,证明核心集选择方法在固定墙钟时间预算下不优于随机采样或全数据训练,其选择成本无法摊销,评估时忽略选择时间会误导结论。
AI 中文摘要
核心集选择从带标签的训练集中挑选一个有代表性的子集,以降低训练成本。然而,它通常通过在固定子集大小下的下游准确率来评估,忽略了选择子集所花费的时间以及每个报告数字背后的训练方案。我们引入了一个端到端的基准测试,该基准标准化了下游训练,并将选择和训练计入同一可审计的墙钟时间预算,涵盖从CIFAR-10到ImageNet-1K的4个数据集、11种选择器、5种比例和3种随机种子,发布了超过1500次运行结果。重复采样的研究表明,预算感知的评估已经偏向随机策略。我们的两项预算研究测试了当每种选择器被赋予其最有利的工作点时,这一结论是否仍然成立。在CIFAR-10和Tiny ImageNet上各8个墙钟时间预算锚点中,没有一个锚点由复杂的选择器获胜:每个获胜者都是类别平衡随机采样、重复随机采样或全数据训练。在ImageNet-1K上的固定预算对决中,用更少的轮次在全数据上训练击败了我们测试的每一种选择策略,同时成本也最低。逐数据集的成本审计显示,选择成本在每个规模上都由一次固定的全数据集扫描主导,因此无法通过选择更小的比例来摊销,其绝对大小也不能从一个数据集外推到另一个数据集。我们进一步量化了选择通过子集复用何时能带来回报,并记录了对一个广泛使用的代码库的9项正确性修复,其中一项使标准的Herding基线提升了近6个百分点。选择时间不是免费的预处理,忽略它的评估衡量了错误的量。
英文摘要
Coreset selection picks a representative subset of the labeled training set to make training cheaper. However, it is usually evaluated by downstream accuracy at a fixed subset size, ignoring both the time spent selecting the subset and the training recipe behind each reported number. We introduce an end-to-end benchmark that standardizes downstream training and charges selection and training to the same auditable wall-clock budget, spanning 4 datasets from CIFAR-10 to ImageNet-1K, 11 selectors, 5 fractions, and 3 seeds, with over 1,500 released runs. Repeated-sampling work has shown that budget-aware evaluation already favors random strategies. Our two budget studies test whether that verdict survives when every selector is granted its most favorable operating point. Across eight wall-clock budget anchors on each of CIFAR-10 and Tiny ImageNet, no anchor is won by a sophisticated selector: every winner is class-balanced random sampling, repeated random sampling, or full-data training. In fixed-budget duels on ImageNet-1K, training on all data for fewer epochs beats every selection strategy we probe while also costing the least. A per-dataset cost audit shows that selection cost is dominated at every scale by a fixed full-dataset scan, so it cannot be amortized away by selecting a smaller fraction, and its absolute size does not extrapolate from one dataset to another. We further quantify when selection does pay back through subset reuse, and document 9 correctness fixes to a widely used codebase, one of which shifts a standard Herding baseline by nearly 6 points. Selection time is not free preprocessing, and an evaluation that ignores it measures the wrong quantity.
Comments23 pages, 7 figures. Benchmark artifacts include per-run result tables, selected indices, and raw timing-audit tables