批量大小还是负样本?内存受限推荐系统训练的选择规则
Batch Size or Negatives? A Selection Rule for Memory-Constrained Recommender Training
- Moscow Independent Research Institute of Artificial Intelligence(莫斯科独立人工智能研究院)
- Intellectual data analysis and predictive modeling institute(智能数据分析与预测建模研究院)
- Applied AI Institute(应用人工智能研究院)
- Risk department(风险部门)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
该研究针对内存受限的推荐系统训练,分析固定内存预算下批量大小与负样本数量的权衡,提出优先扩大批量的选择规则,经实验验证可提升收敛速度与推荐质量。
AI中文摘要:
大规模神经推荐系统通常采用针对全部物品词汇表的softmax交叉熵损失函数进行训练。对于数量庞大的候选物品K,最终分类层会占据主要内存,处理n个样本的批次时需要存储O(nK)个对数几率和梯度。采样softmax通过将损失函数限制在仅k个远小于K的候选负样本上,将内存成本降至O(nk)。然而,在固定预算B = n k的条件下,究竟应优先选择更大的批量还是更多的负样本仍不明确。我们通过分析内存约束下的采样softmax训练来解决该问题。在标准平滑性和方差假设下,理论证据表明最快收敛来自n ~ B、k ~ 1的分配。因此,可执行的规则是在计算约束下尽可能包含更多样本。我们的理论得到受控合成数据及四个真实序列推荐基准(包括MovieLens-20M)的支持。在相同内存约束下,建议的配置相比不平衡替代方案实现了更快的收敛速度和更优的最终推荐质量。这些发现为推荐系统训练期间的内存配置提供了理论和实证基础。代码、可复现材料及所有生成图表的脚本可在该httpsURL获取。
英文摘要:
Large-scale neural recommender systems are typically trained with a softmax cross-entropy objective over the full item vocabulary. For a typical large number of possible items $K$, the final classification layer dominates memory, requiring $O(nK)$ logits and gradients to materialize for a batch of $n$ examples. Sampled softmax reduces this cost by restricting the objective to only $k \ll K$ candidate negative items, resulting in an $O(nk)$ memory. However, for a fixed budget $B = n k$, it remains unclear whether one should prioritize larger batches or the inclusion of more negative items. We address this question by analyzing sampled-softmax training under a fixed memory constraint. Under standard smoothness and variance assumptions, our theoretical evidence suggests that the fastest convergence arises from an $ n \sim B, k \sim 1$ allocation. So, an actionable rule is to include as many objects as possible given computational constraints. Our theory is supported by controlled synthetic and synthetic and four real sequential recommendation benchmarks, including MovieLens-20M. The suggested configuration achieve faster convergence and better final recommendation quality than imbalanced alternatives within the same memory constraint. These findings provide a theoretical and empirical foundation for configuring memory during the training of recommender systems. Code, reproducibility materials, and all scripts for generating figures are available at https://anonymous.4open.science/r/LimitedMemoryRule-BBFB