发表机构
National University of Singapore; Indian Institute of Technology Bombay(新加坡国立大学; 印度理工学院孟买分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对聚类中的coreset构建,提出基于行列式点过程的行列式采样框架,在温和数据分布假设下,将coreset大小对1/ε的依赖指数降至小于2,突破最坏情况ε⁻²下界,并在实验中优于现有方法。
AI 中文摘要
现代机器学习中的海量数据集使得数据缩减成为核心挑战,尤其是在聚类任务中,内存和计算限制要求紧凑而忠实的摘要。标准方法是构建一个 $\epsilon$-coreset:一个小的加权子集,近似保留所有可能中心选择下的聚类成本。对于 $(k,z)$-聚类问题,现有关于 coreset 大小的最坏情况界限基本是紧的,排除了在一般情况下显著更小的 coreset。然而,此类最坏情况实例往往不能代表真实世界数据。在本工作中,我们证明在关于底层数据分布的温和且自然的假设下,显著更小的 coreset 是可能的。我们引入了一种新的相关采样框架,称为行列式采样(determinantal sampling),基于行列式点过程的新应用。利用该框架,我们为 $\mathbb R^d$ 中的 $(k,z)$-聚类获得了一个可高效构造的 $\varepsilon$-coreset,当 $d$ 固定时,其对 $1/\varepsilon$ 的依赖指数严格小于 $2$。这在我们超越最坏情况的假设下改进了最坏情况下的 $\varepsilon^{-2}$ 障碍。据我们所知,这是第一个通过超越最坏情况假设证明性地超越这些下界的结果。最后,我们在合成和真实世界基准数据集上验证了我们的方法,即使在未显式强制执行分析中使用的假设的情况下,它也始终比现有最先进方法实现更小的 coreset。
英文摘要
Massive datasets in modern machine learning have made data reduction a central challenge, particularly for clustering tasks where memory and computational constraints demand compact yet faithful summaries. A standard approach is to construct an \textit{$ε$-coreset}: a small weighted subset that approximately preserves the clustering cost for every plausible choice of centers. For the \textit{$(k,z)$-clustering problem}, existing worst-case bounds on coreset size are essentially tight, ruling out substantially smaller coresets in general. However, such worst-case instances are often unrepresentative of real-world data. In this work, we show that significantly smaller coresets are possible under mild and natural assumptions on the underlying data distribution. We introduce a new correlated sampling framework, called \textit{determinantal sampling}, based on a novel application of determinantal point processes. Using this framework, we obtain an efficiently constructible $\varepsilon$-coreset for $(k,z)$-clustering in $\mathbb R^d$ whose dependence on $1/\varepsilon$ has exponent strictly smaller than $2$ when $d$ is fixed. This improves over the worst-case $\varepsilon^{-2}$ barrier under our beyond-worst-case assumptions. To the best of our knowledge, this is the first result that provably surpasses these lower bounds through beyond-worst-case assumptions. Finally, we validate our approach on synthetic and real-world benchmark datasets, where it consistently achieves smaller coresets than existing state-of-the-art methods, even without explicitly enforcing the assumptions used in the analysis.