大数据K均值聚类的数据原生全局优化
Data-Native Global Optimization for Big Data K-means Clustering
- AI Research Lab, Satbayev University(人工智能研究实验室,萨特巴耶夫大学)
- Laboratory for Analysis and Modeling of Information Processes, Institute of Information and Computational Technologies(信息与计算技术研究所信息过程分析与建模实验室)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对大数据K均值聚类难题,提出Big-means++算法,通过策划输入实现可扩展性与全局搜索质量,利用遍历样本诱导景观、流动在位策略、新抖动机制及竞争多智能体系统等方法,经实验验证其有效性、效率和鲁棒性。
AI中文摘要:
大数据聚类仍然具有挑战性:K均值聚类背后的最小平方和聚类(MSSC)问题是NP难问题,现有方法要么陷入较差的局部最小值,要么需要过高的元启发式混合方法。我们针对任意高维数据,提出了Big-means++算法,通过精心策划输入来实现大数据MSSC优化的可扩展性和全局搜索质量。它将局部K均值细化编排成大数据聚类的数据原生全局搜索。Big-means++遍历样本诱导的替代景观,通过流动在位策略在经验景观中传播质心状态,新的抖动机制改变样本大小,竞争多智能体系统异步探索独立采样景观,自动收敛检测在达到高质量解决方案后停止每个智能体。在22个数据集上与11种竞争算法进行的实验证明了Big-means++的有效性、效率和鲁棒性。
英文摘要:
Big data clustering remains challenging: the Minimum Sum-of-Squares Clustering (MSSC) problem underlying K-means is NP-hard, and existing methods either reach poor local minima or require prohibitive metaheuristic hybrids. We target arbitrarily tall data: a fixed feature space may contain arbitrarily many, possibly infinitely many, observations, while the algorithm accesses only finite random samples. We propose Big-means++, an algorithm achieving scalability and global-search quality by curating inputs to MSSC optimization on big data. It orchestrates local K-means refinements into a data-native global search for big data clustering. Rather than optimizing the full-data MSSC objective, Big-means++ traverses sample-induced surrogate landscapes. Each sample defines a distinct empirical MSSC approximation with a perturbed local-optimum structure, turning sample-to-sample variation into a global-search mechanism. Unlike Big-means, a flowing-incumbent strategy propagates centroid state across empirical landscapes through K-means refinements on fresh samples without rollback to a best-so-far solution. This increases mobility and favors stable, high-quality configurations across approximations of the full-data structure. A new shaking mechanism varies sample size geometrically, broadening the surrogate landscapes explored across resolution scales, accounting for cluster imbalance, and improving solution quality. A competitive multi-agent system asynchronously explores independent sampled landscapes, transforming diverse stochastic trajectories into collective search intelligence. Automatic convergence detection stops each agent after attaining a high-quality solution but before further search risks degrading it, while providing a universal speed-quality control. Experiments on 22 datasets against 11 competing algorithms demonstrate the effectiveness, efficiency, and robustness of Big-means++.