arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.07432cs.LGcs.AI

重新审视核学习问题中的稀疏化方法

Revisiting Thinning Methods for Kernel Learning Problems

  • Universidad Autónoma de Madrid(马德里自治大学)

机构由 AI 辅助整理,请以论文原文为准。

Blanca Cano-Camarero, Yago R. Aguado-Carrillo-de-Albornoz, Ángela Fernández-Pascual, José R. Dorronsoro

AI总结:

本文提出反向核拾取与灵活核稀疏化两种改进的核方法数据缩减算法,在保持预测性能的同时显著提升训练效率,并系统比较了不同场景下的适用性。

AI中文摘要:

核方法因其强大的理论保证和实证性能而被广泛使用。然而,其高昂的计算成本限制了其在大规模数据集上的适用性。为了解决这一缺点,几种方法使用最大均值差异来构造代表性子集,以在再生核希尔伯特空间中保留完整数据集的属性。我们引入了反向核拾取(Backward Kernel Herding),这是一种通过迭代地从数据集中移除点来解决该问题的算法,在缩减规模小于数据集一半的现实场景中,其性能可与当前最先进的方法相媲美,同时加速了子采样过程。此外,我们克服了核稀疏化(Kernel Thinning)的一个局限性,提出了一种扩展方法,使得能够构造任意大小的子集,而非仅限于连续的减半操作。最后,我们针对最相关的核学习过程——高斯过程和核支持向量机——进行了广泛的实验比较。结果表明,反向核拾取始终以最有利的训练时间效率取得具有竞争力的性能,而所提出的灵活核稀疏化(Flexible Kernel Thinning)则经常取得最佳的预测性能。这些优势在中等压缩比下尤为明显,凸显了将监督信息纳入稀疏化过程的好处。在内存消耗方面,灵活核稀疏化同样具有竞争力,而反向核拾取在计算效率为首要目标时仍是一种替代方案。总体而言,没有任何单一方法在所有场景中占主导地位,这强调了根据预测性能、训练成本和内存需求之间的期望权衡来选择缩减策略的重要性。

英文摘要:

Kernel methods are widely used because of their strong theoretical guarantees and empirical performance. However, their high computational cost limits their applicability to large-scale datasets. To address this shortcoming, several approaches use Maximum Mean Discrepancy to construct representative subsets that preserve the properties of the full dataset in a Reproducing Kernel Hilbert Space. We introduce Backward Kernel Herding, an algorithm that addresses this problem by iteratively removing points from the dataset, achieving results comparable to current state-of-the-art approaches while accelerating the subsampling process in realistic scenarios where the reduced size is less than half of the dataset. Moreover, we overcome a limitation of Kernel Thinning by proposing an extension that enables the construction of subsets of arbitrary size rather that restricting to successive halvings. Finally, we conduct an extensive experimental comparison focusing on the most relevant kernel learning procedures: Gaussian Processes and Kernel Support Vector Machines. The results show that Backward Kernel Herding consistently achieves competitive performance with the most favorable training-time efficiency, while the proposed Flexible Kernel Thinning frequently achieves the best predictive performance. These gains become especially pronounced for moderate compression ratios, highlighting the benefits of incorporating supervised information into the thinning process. In terms of memory consumption, Flexible Kernel Thinning is also competitive, whereas Backward Kernel Herding remains an alternative when computational efficiency is the primary objective. Overall, no single method dominates across all scenarios, underscoring the importance of selecting the reduction strategy according to the desired trade-off between predictive performance, training cost, and memory requirements.

↑