基于分位数的损失过滤用于抗离群值随机梯度下降
Quantile-based Loss Filtering for Outlier-Robust Stochastic Gradient Descent
- Harvey Mudd College(哈维穆德学院)
- University of California, Irvine(加州大学尔湾分校)
- University of Oxford(牛津大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本研究提出分位数-k损失SGD(QkL-SGD)框架,通过采样损失并选择较低分位数索引更新,在凸假设下证明线性收敛,实验显示中间分位数优于标准SGD和min-k-loss,兼具鲁棒性与信息量。
AI中文摘要:
我们研究了在有限和优化中基于损失的过滤方法,其中一部分分量函数的梯度可能高度不可靠。受最小损失随机梯度下降(min-k-loss)和针对损坏线性系统的分位数方法的启发,我们提出并分析了一个通用的损失过滤框架——分位数-k损失随机梯度下降(QkL-SGD),该框架在每次迭代中采样k个分量损失,并使用从较低经验q分位数中均匀选择的索引进行更新。我们在标准凸性假设下证明了这类方法的线性收敛性,要求样本大小与损坏数量成比例,并满足子集强凸性阈值。对于无法或不希望进行足够大采样的情形,我们给出了一个补充的小样本概率分析,该分析涵盖任意样本大小k,且收敛行为取决于选择离群值的概率以及所选良好步骤的曲率。在多项式回归、正则化逻辑回归和正则化合页损失上的实验表明,中间分位数通常优于标准SGD和min-k-loss SGD。特别是,min-k经常因重复选择几乎已解决的分量而停滞,而中间分位数则保持鲁棒性并产生更具信息量的更新。
英文摘要:
We study loss-based filtering for finite-sum optimization with a subset of corrupted component functions whose gradients may be highly unreliable. Motivated by minimum-loss-based SGD (min-$k$-loss) and quantile-based methods for corrupted linear systems, we propose and analyze a general loss-filtering framework -- Quantile-\(k\)-Loss SGD (Q\(k\)L-SGD) -- that samples \(k\) component losses at each iteration and updates using an index chosen uniformly from the lower empirical \(q\)-quantile. We prove linear convergence of this family of methods under standard convexity assumptions, requiring the sample size to scale with the number of corruptions and a subset strong-convexity threshold. For the cases when large enough sampling is impossible or undesirable, we give a complementary small-sample probabilistic analysis that covers any sample size $k$ and the convergence behavior depends on the probability of selecting an outlier and on the curvature of the selected good step. Experiments on polynomial regression, regularized logistic regression, and regularized hinge loss show that intermediate quantiles often outperform both standard SGD and min-\(k\)-loss SGD. In particular, min-\(k\) often stalls by repeatedly selecting nearly solved components, while intermediate quantiles retain robustness and produce more informative updates.