Wasserstein滤波:一种用于鲁棒分布学习的样本选择方法
Wasserstein Filtering: A Sample Selection Method for Robust Distribution Learning
浏览论文内容
中文总结 AI 辅助
针对存在部分样本被污染的数据集,提出Wasserstein滤波样本选择框架,引入三种算法,在FELP污染模型下证明其估计量的极小极大最优性,经实验验证其为实用的模型无关预处理工具,兼具良好异常检测性能与生成建模下游效益。
中文摘要 AI 辅助
给定存在部分样本被污染的数据集,我们的目标是恢复潜在的干净总体分布。为此,我们提出Wasserstein滤波(WF),这是一种新颖的样本选择框架,它丢弃一部分可疑样本,并利用剩余数据的经验测度估计目标分布。核心思路是选择一个样本子集,其经验分布与完全污染的经验分布的Wasserstein距离最大化,从而优先隔离和移除具有几何影响力的异常值。为使该优化问题具有计算可处理性,我们引入三种算法:边际筛选方案SinkMarg,以及两种联合优化算法SinkWF和SlicedWF,分别利用熵最优传输和切片Wasserstein近似。理论方面,我们提出远距排除与局部投影(FELP)污染模型,该模型刻画由分离良好的异常值和局部不可区分的扰动组成的污染。在该模型下,我们证明WF估计量在协方差有界的分布族上达到极小极大最优性。在合成数据集、基准异常检测套件以及基于扩散模型的鲁棒生成学习上开展的大量数值实验表明,WF是一种实用性强、与模型无关的预处理工具,它具备有竞争力的异常值检测性能,并在严重污染下为生成建模提供显著的下游效益。
英文摘要
Given a dataset where a portion of the samples are contaminated, our goal is to recover the underlying clean population distribution. To this end, we propose Wasserstein Filtering (WF), a novel sample selection framework that discards a fraction of suspicious samples and estimates the target distribution using the empirical measure of the remaining data. The core insight is to select a subset of samples whose empirical distribution maximizes its Wasserstein distance to the fully contaminated empirical distribution, thereby preferentially isolating and removing geometrically influential outliers. To render this optimization computationally tractable, we introduce three algorithms: a marginal screening scheme, SinkMarg, and two joint optimization algorithms, SinkWF and SlicedWF, leveraging entropic optimal transport and sliced Wasserstein approximations, respectively. On the theoretical front, we introduce the Far Exclusion and Local Projection (FELP) contamination model, which characterizes corruptions consisting of well-separated outliers and locally indistinguishable perturbations. Under this model, we prove that the WF estimator achieves minimax optimality over distribution families with bounded covariance. Extensive numerical experiments on synthetic datasets, benchmark anomaly detection suites, and robust generative learning with diffusion models demonstrate that WF serves as a highly practical, model-agnostic preprocessing tool. It delivers competitive outlier detection performance and provides substantial downstream benefits for generative modeling under heavy contamination.
发表机构
- The Hong Kong Polytechnic University(香港理工大学)
- Fudan University(复旦大学)
机构由 AI 辅助整理,请以论文原文为准。