发表机构
Spotify; SJTU Paris Elite Institute of Technology(Spotify; 上海交通大学巴黎卓越工程师学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对两级softmax采样因忽略簇大小不平衡和簇内离散度而产生的偏差,提出S-2LS和SD-2LS两种校正方法,在五个大规模数据集上验证了改进的采样性能,建议替代标准2LS。
AI 中文摘要
从softmax分布中采样是机器学习中的一项基本操作,但其在项目数量上的线性复杂度使得在大规模场景下进行精确采样变得不切实际。两级softmax(2LS)采样是一种流行的替代方案,能够实现次线性时间采样。假设项目被划分为若干簇,2LS首先采样一个簇,然后在该簇内采样一个项目。在本文中,我们表明,尽管2LS具有优势,但它引入了系统性的、不期望的采样偏差,这些偏差源于对簇的错误加权,忽略了簇大小不平衡和簇内相似性离散度。我们提出了两种采样方法,即大小校正的2LS(S-2LS)和大小与离散度校正的2LS(SD-2LS),它们纠正了这些偏差,并提供了可证明更好的softmax近似,且计算开销可忽略不计甚至不存在。在五个大规模数据集上的深入实验验证了我们方法改进的采样特性。我们建议在未来的工作中一致使用它们来代替标准的2LS。
英文摘要
Sampling from a softmax distribution is a fundamental operation in machine learning, but its linear complexity in the number of items makes exact sampling impractical at scale. Two-level softmax (2LS) sampling is a popular alternative enabling sublinear-time sampling. Assuming items are partitioned into clusters, 2LS first samples a cluster and then an item within it. In this paper, we show that, despite its advantages, 2LS introduces systematic and undesirable sampling biases, which arise from misweighting clusters by ignoring both cluster size imbalance and intra-cluster similarity dispersion. We propose two sampling methods, Size-Corrected 2LS (S-2LS) and Size- and Dispersion-Corrected 2LS (SD-2LS), which correct these biases and provide provably better softmax approximations with negligible to non-existent computational overhead. In-depth experiments on five large-scale datasets validate the improved sampling properties of our methods. We recommend their consistent use in place of standard 2LS in future work.
CommentsNeurIPS 2026