AI 中文总结
本文提出贝叶斯域加权方法,引入Gamma先验从狄利克雷分布推断权重,实现稳定高效的域权重学习,以更少数据识别最优数据混合,适用于大规模应用。
AI 中文摘要
大语言模型(LLMs)的性能从根本上受多领域预训练数据的分布构成影响。早期模型普遍采用手动启发式方法,但随着数据复杂性提升,这类方法越来越无法捕捉域间的复杂协同作用。为解决该问题,主流方法是拟合域权重与对应验证损失间的代理函数,再寻找最优域权重以最小化验证损失。这些方法依赖强结构假设,如秩不变性或缩放定律,而这些假设常被违反,导致不可忽视的估计偏差。另一种有前景的方法是直接优化数据的加权方案,但存在优化轨迹不稳定、计算开销过高的问题,限制了其搜索更优域权重配置的潜力。本文提出一种贝叶斯域加权方法,通过引入从观测中学习到的Gamma先验信息,从狄利克雷分布中推断权重。实验结果表明,该方法可实现稳定高效的域权重学习,且与基于搜索的函数拟合方法相比,消耗的数据量显著更少,能识别最优混合方案,为大规模应用重新激活了基于优化的域加权方法。
英文摘要
The performance of Large Language Models (LLMs) is fundamentally influenced by the distributional composition of multi-domain pre-training data. While manual heuristics were prevalent in early models, they increasingly fail to capture the intricate synergies between domains as data complexity grows. To overcome the issue, a dominant approach seeks to fit a proxy function mapping between domain weights and their corresponding validation losses, and then find the optimal domain weights to minimize validation losses. These methods rely on strong structural assumptions, such as rank invariance or scaling laws, which are often violated, resulting in non-negligible estimation bias. A promising approach is to directly optimize the weighting scheme from data. However, it suffers from unstable optimization trajectory and prohibitive computational overhead, limiting its potential to search better domain weights configurations. This paper presents a Bayesian domain weighting method to infer the weights from a Dirichlet distribution via introducing Gamma prior information learned from observations. Experimental results demonstrate that proposed method could achieve stable and efficient domain weights learning, and identifies optimal mixtures while consuming substantially less data than search-based function-fitting methods, revitalizing optimization-based domain weighting for large-scale applications.