Hamming空间中计算核心集与k-中位数聚类的流算法
Streaming algorithms for computing coresets and $k$-median clustering in the Hamming space
AI总结:
针对Hamming空间中W[1]-困难的连续k-中位数聚类问题,提出首个带(1+ε)近似比的FPT流算法,空间复杂度低,还给出该问题首个ε-核心集的流算法。
AI中文摘要:
聚类是数据分析中最基础的工具之一,可通过少量代表性点对大型数据集进行概括。给定度量空间$(\t{X}, \t{d})$及该空间中$n$个点的集合$S$,连续k-中位数聚类问题旨在找到包含$k$个点的集合$C$,使目标函数$\textstyle\boldsymbol{\text{sum}}_{s \boldsymbol{\text{in}} S} \t{d}(s,C)$最小。当$\t{X} = \boldsymbol{\text{Σ}}^\boldsymbol{\text{ℓ}}$为长度$\boldsymbol{\text{ℓ}}$的字符串集合,且$\t{d}$为Hamming距离时,该问题在以$k$为参数时被证明是W[1]-困难的。本研究提出了该问题的首个$(1+\boldsymbol{\text{ε}})$近似算法,其FPT运行时间为$2^{\boldsymbol{\text{poly}}(\boldsymbol{\text{ε}}^{-1},k)} \boldsymbol{\text{·}} n\boldsymbol{\text{ℓ}} \boldsymbol{\text{polylog}} n$。该算法的另一特性是可在流场景下实现,仅需$\tilde{\boldsymbol{\text{O}}}_\boldsymbol{\text{ε}}(\boldsymbol{\text{ℓ}}k + k^2)$的空间。作为一项具有独立研究价值的辅助工具,本文还展示了首个针对Hamming空间下连续k-中位数聚类计算$\boldsymbol{\text{ε}}$-核心集的流算法。
英文摘要:
Clustering is one of the most fundamental tools in data analysis, allowing large datasets to be summarized by a small number of representative points. Given a metric space $(\mathcal{X}, \mathbb{d})$ and a set $S$ of $n$ points in this space, the continuous $k$-median clustering problem asks to find a set $C$ of $k$ points that minimizes the objective function $\sum_{s\in S} \mathbb{d}(s,C)$. When $\mathcal{X} = Σ^\ell$ is the set of strings of length $\ell$ and $\mathbb{d}$ is the Hamming distance, the continuous $k$-median clustering problem is known to be W[1]-hard when parameterized by $k$. In this work, we present the first $(1+\varepsilon)$-approximation algorithm for this problem with FPT runtime $2^{\mathrm{poly}(\varepsilon^{-1},k)} \cdot n\ell \mathrm{polylog} \; n$. An additional feature of the algorithm is that it can be implemented in streaming, requiring only $\tilde{O}_\varepsilon(\ell k + k^2)$ space. As an auxiliary tool of independent interest, we show the first streaming algorithm for computing an $\varepsilon$-coreset for continuous $k$-median clustering under the Hamming