聚类的新视角:混合范数模型及其通过渐进整数规划的求解
A New Perspective on Clustering: A Mixed-norm Model and its Solution by Progressive Integer Programming
浏览论文内容
中文总结 AI 辅助
本文提出一种混合范数聚类模型,通过渐进整数规划方法高效求解,并证明其在多种污染场景下优于经典聚类方法。
中文摘要 AI 辅助
本文在经典 $K$-means 和 $K$-medians 模型的基础上,引入了一个 $\ell_{p,q}$ 混合范数聚类模型,其中质心更新和簇分配分别受 $\ell_p$ 和 $\ell_q$ 范数约束。该模型被表述为一个混合整数规划(MIP),带有描述最近中心分配的 Heaviside 复合约束。该框架在 $p=q=2$ 和 $p=q=1$ 时分别恢复 $K$-means 和 $K$-medians 模型,并在 $p\ne q$ 时产生新模型。为了解决计算挑战,我们开发了一种渐进整数规划(PIP)方法,该方法自适应地固定置信分配并求解受限的混合整数子问题。对于 $q=1$,我们在受限子问题中开发了凸约束的差之凸内逼近,可以计算全局解。重要的是,我们建立了混合范数聚类问题的局部极小点与强中心局部极小点之间的联系,以及在特定假设下与受限子问题的全局最优解之间的联系。这种联系为非凸混合范数聚类模型的局部极小点提供了实用的证书。我们进一步开发了为 $q=1$ 构建自适应固定集和工作集的技术。大量数值实验证明了混合范数聚类模型的优越性能和 PIP 在求解 MIP 模型(否则可能难以处理)方面的效率。特别是,$\ell_{2,1}$ 混合范数聚类模型在坐标稀疏、均值平衡污染下有效,而 $\ell_{1,2}$ 模型在密集坐标柯西污染下更受青睐。数值结果还表明,PIP 可以逃离较差的交替解并获得显著更好的可行聚类,同时在未发现改进时保持强热启动。
英文摘要
Extending the classical $K$-means and $K$-medians models, this paper introduces an $\ell_{p,q}$ mixed-norm clustering model where the centroid updates and cluster assignments are under the $\ell_p$ and $\ell_q$ norms, respectively. The model is formulated as a mixed-integer program (MIP) with Heaviside composite constraints that describe the nearest-center assignments. The framework recovers $K$-means and $K$-medians when $p=q=2$ and $p=q=1$, respectively, and yields new models when $p\ne q$. To address the computational challenges, we develop a progressive integer programming (PIP) method that adaptively fixes confident assignments and solves restricted mixed-integer subproblems. For $q=1$, we develop a convex inner approximation of the difference-of-convex constraints in the restricted subproblems, for which a global solution can be computed. Importantly, we establish the connection between the local minimizer and the strong center-local minimizer of the mixed-norm clustering problem and the global optimal solution of the restricted subproblems under certain assumptions. This connection provides a practical certificate of a local minimizer of the nonconvex mixed-norm clustering model. We further develop techniques for constructing adaptive fixing sets and working sets for $q=1$. Extensive numerical experiments demonstrate the superior performance of the mixed-norm clustering model and the efficiency of PIP for solving the MIP model, which may be intractable otherwise. In particular, the $\ell_{2,1}$ mixed-norm clustering model is effective under coordinate-sparse, mean-balanced contamination, whereas the $\ell_{1,2}$ model is preferred under dense coordinatewise Cauchy contamination. The numerical results also show that PIP can escape poor alternating solutions and obtain substantially better feasible clustering, while preserving strong warm starts when no improvement is found.
发表机构
- Tsinghua University(清华大学)
- The Daniel J. Epstein Department of Industrial and Systems Engineering, University of Southern California(南加州大学丹尼尔J·艾普斯坦工业与系统工程系)
- H. Milton Stewart School of Industrial and Systems Engineering, Georgia Institute of Technology(佐治亚理工学院H·米尔顿·斯图尔特工业与系统工程学院)
- Department of Applied Mathematics, Hong Kong Polytechnic University(香港理工大学应用数学系)
机构由 AI 辅助整理,请以论文原文为准。