发表机构
Columbia University; University of California, Santa Cruz(哥伦比亚大学; 加州大学圣克鲁兹分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出DUCB-OGD算法,结合折扣上置信界采样与在线梯度下降,解决LLM后训练中组分布鲁棒的动态极小极大遗憾问题,实现最优遗憾界并提升最差组鲁棒性。
AI 中文摘要
现代大语言模型(LLM)训练日益依赖于涵盖不同领域、任务、偏好分布和难度级别的异构数据源。我们研究了在仅瞬时小批量强盗反馈下,针对组分布鲁棒的LLM后训练的动态极小极大遗憾问题。该框架将训练视为一个双人采样器-优化器过程:采样器利用强盗反馈自适应地在数据源之间进行选择,而优化器则使用来自所选源的随机梯度更新模型参数。我们聚焦于实际受限的设置,其中源损失随模型训练而变化,但历史数据不被重新评估,这要求采样器从过时的部分反馈中跟踪瞬时最差源。我们提出了DUCB-OGD,一种简单且可扩展的算法,将折扣上置信界(DUCB)采样器与在线梯度下降(OGD)优化器相结合。采样器维护基于折扣有效样本量的指数移动平均损失估计和置信半径,避免了对过去数据的昂贵重新评估或对标准训练流程的侵入性更改。对于$K$个数据源和$T$个训练步骤,我们证明DUCB-OGD实现了$\tilde{O}(K^{1/4}T^{3/4})$的动态极小极大遗憾,这在我们的反馈模型下对于非折扣目标是最优的(达到对数因子)。在监督微调、偏好优化和强化学习中的广泛实验表明,DUCB-OGD能无缝集成到现代LLM训练流程中,并以可忽略的计算开销相比标准采样基线提高了最差组的鲁棒性。
英文摘要
Modern LLM training increasingly relies on heterogeneous data sources spanning different domains, tasks, preference distributions, and difficulty levels. We study dynamic minimax regret for group-distributionally robust LLM post-training under instantaneous mini-batch-only bandit feedback. The framework views the training as a two-player sampler-optimizer process: a sampler adaptively selects among data sources using bandit feedback, while an optimizer updates the model parameters using stochastic gradients from the selected source. We focus on the practically restrictive setting where source losses evolve with model training but historical data are not re-evaluated, requiring the sampler to track instantaneous worst-sources from stale partial feedback. We propose DUCB-OGD, a simple and scalable algorithm that couples a Discounted Upper-Confidence-Bound sampler with an Online Gradient Descent optimizer. The sampler maintains exponential moving average loss estimates and confidence radii based on discounted effective sample sizes, avoiding costly re-evaluation of past data or intrusive changes to standard training pipelines. For $K$ data sources and $T$ training steps, we prove that DUCB-OGD achieves a dynamic minimax regret of $\tilde{O}(K^{1/4}T^{3/4})$, which is optimal up to logarithmic factors for the undiscounted objective under our feedback model. Extensive experiments across supervised fine-tuning, preference optimization, and reinforcement learning show that DUCB-OGD integrates seamlessly into modern LLM training pipelines and improves worst-group robustness with negligible computational overhead compared with standard sampling baselines.