在线差分隐私一致性聚类
Online Differentially Private Consistent Clustering
AI总结:
该研究针对在线流环境下的差分隐私聚类问题,提出通用归约方法,使算法在近似比、空间与时间性能上优于或匹配现有方法,且继承了非私有聚类算法的一致性属性。
AI中文摘要:
我们研究在线流环境下的差分隐私(DP)k-均值与k-中位数聚类问题。该模型中,点按顺序到达,每一步需输出k个中心的集合,以优化当前已观测到所有点的聚类目标。我们提出一种通用归约方法,将敏感输入流转换为私有流,该私有流是输入流的半核心集。这意味着任何(非私有)在线聚类算法作为后处理步骤运行时,均可针对原始聚类目标实现良好的效用。我们的算法在近似比、空间使用量和运行时间上与现有算法[Epasto等人,2026;Dupré la Tour等人,2024]相当或更优。该归约方法的关键特性是继承了底层非私有聚类算法的理想属性,例如一致性[Lattanzi和Vassilvitskii,2017]——这是此前差分隐私算法未满足的属性。
英文摘要:
We study differentially private (DP) $k$-means and $k$-median clustering in the online streaming setting. In this model, points arrive sequentially, and at each time step, we need to output a set of $k$ centers that optimizes the clustering objective for all points seen so far. We give a generic reduction that transforms the (sensitive) input stream into a private stream, which is a semi-coreset of the input stream. This implies that any (non-private) online clustering algorithm, run as a post-processing step, can achieve good utility for the original clustering objective. Our algorithm matches or improves upon the approximation ratio, space usage, and running time of existing algorithms [Epasto et al., 2026, Dupré la Tour et al., 2024]. A key aspect of our reduction is that it inherits desirable properties of the underlying non-private clustering algorithm, such as consistency [Lattanzi and Vassilvitskii, 2017]--a property not satisfied by previous DP algorithms.