arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于第一性原理的数据集剪枝:一种无标签的线性规划方法

Dataset Pruning from First Principles: A Label-Free Linear Programming Approach

Rodrigo Schuller, Francisco Ganacim

arXiv 2610.10347首次发表:更新:

发表机构

Instituto de Matemática Pura e Aplicada (IMPA)(纯数学与应用数学研究所(IMPA))

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出一种基于方差最小化的无标签线性规划数据集剪枝方法,通过顶点游走优化无偏子集选择,在多个基准上匹配或超越均匀采样,并降低随机梯度方差。

AI 中文摘要

数据集剪枝在保持模型性能的同时,将大型训练集缩减为具有代表性的子集。现有的基于几何的方法通常假设嵌入空间中邻近的点具有相似的性质。我们不采用这一假设,而是通过将无偏子集选择重新表述为方差最小化问题来推导几何选择准则。无偏性确保未加权的子集平均值在期望上恢复完整数据集的平均值,包括在固定模型参数下的损失和梯度。具体而言,我们将一族无偏子集选择算法表征为一个高维多面体。在此背景下,最小化期望采样方差是一个线性目标。在刚性运动上平均的采样方差差异,允许闭式成对表达式。由于多面体具有高维性,直接应用标准线性规划是不切实际的。我们转而利用这些表达式构建一种高效的顶点游走方法,在保持无偏性的同时优化方差目标的近似,从而得到一种在选择过程中既不需要标签也不需要模型训练的方法。在CIFAR-10、MNIST和CelebA基准测试中,我们的方法在每个评估预算下的平均测试准确率上匹配或超过均匀采样,并在若干设置中优于竞争的几何方法,尤其是在小选择预算下。除数据集剪枝外,同一框架通过增加小批量内的多样性同时保持批量大小不变,降低了随机梯度方差。

英文摘要

Dataset pruning reduces a large training set to a representative subset while preserving model performance. Existing geometry-based methods typically assume that nearby points in embedding space share similar properties. Rather than imposing this assumption, we derive geometric selection criteria by reformulating unbiased subset selection as a variance minimization problem. Unbiasedness ensures that unweighted subset averages recover full-dataset averages in expectation, including losses and gradients at fixed model parameters. Specifically, we characterize a family of unbiased subset selection algorithms as a high-dimensional polytope. In this context, minimizing the expected sampling variance is a linear objective. Differences in sampling variance, averaged over rigid motions, admit closed-form pairwise expressions. Because the polytope has high dimension, directly applying standard linear programming is impractical. We instead use these expressions to construct an efficient vertex walk that optimizes an approximation of the variance objective while preserving unbiasedness, yielding a method that requires neither labels nor model training during selection. Across CIFAR-10, MNIST, and CelebA benchmarks, our method matches or exceeds uniform sampling in mean test accuracy at every evaluated budget and outperforms competing geometric methods in several settings, particularly at small selection budgets. Beyond dataset pruning, the same framework reduces stochastic-gradient variance by increasing diversity within mini-batches while keeping the batch size unchanged.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑