数据稀缺与模型稀疏:混合专家模型对重复数据的过拟合更严重
Data Scarcity and Model Sparsity: Mixtures-of-Experts Overfit More to Repeated Data
- Stanford University(斯坦福大学)
- Paul G. Allen School of Computer Science, University of Washington(华盛顿大学保罗·G·艾伦计算机科学学院)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文研究数据重复对混合专家模型的影响,发现其比密集模型过拟合更严重,且随稀疏性加剧;强掩码正则化可缓解,但无法完全替代唯一数据。
AI中文摘要:
随着人类书写文本供应的枯竭,重复使用语言模型训练数据已成为标准做法。先前的工作研究了密集激活Transformer中的数据重复问题,但数据重复对近期占主导地位的稀疏架构(如混合专家模型(MoE))的影响在很大程度上仍未得到探索,尽管这些架构具有更高的计算效率。我们在单领域和多领域数据混合中,以及在不同MoE设置(包括专家数量和粒度)下,变化数据重复率。我们一致发现,对于从80M到1B激活参数(总计8.5B参数)的模型,MoE在数据重复下性能下降更快。这种效应随稀疏性增加而增强,由总参数而非激活参数决定。虽然80M的密集模型可以重复数据超过8倍而性能下降最小,但MoE在4倍时就开始受损,并迅速恶化,在32倍后失去其在全唯一数据设置中的性能优势,表现不如密集模型。我们尝试了现有的正则化方法作为潜在补救措施。我们发现某些方法(如dropout)可以缓解过拟合。特别是,在强掩码正则化下,即使数据重复超过64次,MoE也能优于密集模型。然而,没有任何方法能完全匹配全唯一训练数据的性能。最后,我们分析了与高重复率下MoE过拟合相关的内部机制,发现MoE路由在训练早期普遍稳定,且专家专业化与对重复数据的过拟合相关。总之,我们的工作解决了稀疏性与数据重复之间的不利交互:我们提供了过拟合核心机制及其潜在补救的证据,并提出了未来通过破坏记忆模式来减少模型参数过度专业化的有前景的方法途径。
英文摘要:
As the supply of human-written text is exhausted, it has become standard practice to repeat language model training data. Prior work has studied data repetition for densely activated Transformers, but the effects of data repetition remains largely unexplored for recently dominant sparse architectures such as Mixture-of-Experts (MoE), despite their increased compute efficiency. We vary data repetition rates across single- and multi-domain data mixes, and across MoE settings, including expert count and granularity. We consistently find, for models ranging from 80M to 1B active (8.5B total) parameters, that MoEs degrade more rapidly under data repetition. This effect increases with sparsity, dictated by total rather than active parameters. While 80M dense models can repeat data over 8x with minimal degradation, MoEs instead begin to suffer at 4x, and deteriorate rapidly, ceding their performance benefits in all-unique data settings to underperform dense models after 32x. We experiment with existing regularization methods as a potential remedy. We find that some methods, such as dropout, can mitigate overfitting. In particular, with strong masking-based regularization, MoEs are able to outperform dense models even when data is repeated more than 64 times. However, no method fully matches the performance of all-unique training data. Finally, we analyze internal mechanisms correlated with MoE overfitting in high repetition regimes, and find that MoE routing universally stabilizes early in training, and that expert specialization correlates with overfitting to repeated data. In sum, our work addresses the adverse interactions between sparsity and data repetition: we present evidence for the core mechanisms of overfitting and its potential remediation, and suggest promising avenues for future methods to reduce over-specialization in model parameters by disrupting memorization patterns.