arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

鲁棒最大最小多样化的流式算法

Streaming algorithms for robust max-min diversification

Andrea Pietracaprina, Geppino Pucci, Stefano Zanon

arXiv 2610.01456首次发表:更新:

发表机构

University of Padova(帕多瓦大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对鲁棒最大最小多样化问题,提出一种确定性基于核心集的流式算法,在内部点-离群点分离假设下返回恰好 $k$ 个内部点,达到 $(2+\varepsilon)$ 近似,内存和更新代价与 $n$ 无关。

AI 中文摘要

给定度量空间中的一组 $n$ 个点 $X$ 和一个整数 $k$,最大最小多样化旨在选择 $X$ 中的 $k$ 个点,最大化它们之间的最小成对距离。然而,该目标函数极易受到噪声点的影响。在 [Amagata, AAAI23] 中,提出了一种鲁棒公式,通过排除包含 $z$ 个离群点(定义为 $X$ 中具有最大最近邻距离的 $z$ 个点)的任何解来解决此脆弱性。该论文还提出了一种基于核心集的流式算法,该算法基于合适的内部点-离群点分离假设。然而,我们发现了 [Amagata, AAAI23] 算法的三个缺点:其核心集构建需要对 $X$ 进行离线计算,这需要线性于 $n$ 的内存,与流处理的典型目标形成鲜明对比;用于从核心集提取解的单遍过程可能返回少于 $k$ 个点(因此产生不可行解),因为它永久丢弃了距离当前解太远的点;其离群点排除保证仅是概率性的,并且随着核心集大小的缩小而减弱。相比之下,我们提出了一种确定性的基于核心集的算法,在自然的内部点-离群点分离假设(类似于 [Amagata, AAAI23] 中使用的假设)下,该算法返回恰好 $k$ 个内部点,它们是 $(2+\varepsilon)$ 近似解,对于任何 $\varepsilon>0$,因此仅比最佳多项式时间顺序近似高出 $\varepsilon$,即使没有离群点也是如此。其单遍流式实现自适应地适应数据集的双倍维度 $D$,并且对于广泛的 $k$、$z$、$\varepsilon$ 和 $D$ 值,它使用的内存与 $n$ 无关。对于足够长的流,其摊还更新时间与核心集大小成正比,因此也与 $n$ 无关。

英文摘要

Given a set of $n$ points $X$ in a metric space and an integer $k$, max-min diversification aims to select $k$ points of $X$ maximizing their minimum pairwise distance. This objective function is however highly vulnerable to noisy points. In[Amagata, AAAI23], a robust formulation is proposed which addresses this vulnerability by excluding solutions containing any of $z$ outliers, defined as the $z$ points in $X$ with the largest nearest-neighbor distances. That paper also presents a coreset-based streaming algorithm for the new formulation, based on a suitable inlier-outlier separation assumption. However, we identify three shortcomings in the algorithm by [Amagata, AAAI23]: its coreset construction requires an offline computation over $X$, which needs memory linear in $n$, in stark contrast with the typical goals of stream processing; the one-pass procedure used to extract the solution from the coreset may return fewer than $k$ points (hence, an unfeasible solution) because it permanently discards points too far from the current solution; and its outlier-exclusion guarantee is only probabilistic and weakens as the coreset size shrinks. In contrast, we present a deterministic coreset-based algorithm that, under a natural inlier-outlier separation assumption (similar to the one used in [Amagata, AAAI23]), returns exactly $k$ inliers which are a $(2+\varepsilon)$-approximate solution, for any $\varepsilon>0$, thus only $\varepsilon$ above the best polynomial-time sequential approximation, even without outliers. Its one-pass streaming implementation adapts obliviously to the dataset's doubling dimension $D$ and, for wide ranges of $k$, $z$, $\varepsilon$, and $D$, it uses memory independent of $n$. For sufficiently long streams, its amortized update time is proportional to the coreset size, thus also independent of $n$.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑