发表机构
University of Edinburgh; National Technical University of Athens; Athena/Archimedes Research Centre(爱丁堡大学; 雅典国立技术大学; 雅典娜/阿基米德研究中心)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文针对非光滑、超线性梯度增长且非凸的目标分布采样问题,提出SG-TULA算法,推导其非渐近收敛界,验证假设并用于GPT-2系列LLM预训练,效果优于AdamW等。
AI 中文摘要
我们研究从同时具有非光滑、超线性梯度增长和非凸特性的势函数的目标分布中采样的问题。我们引入次梯度驯服未校正朗之万算法(SG-TULA),这是一种朗之万扩散的离散化方法,直接作用于次梯度,无需依赖计算成本高昂的平滑过程。为处理超线性区域,采用驯服技术生成稳定的显式方案。我们在Wasserstein-2距离中推导了非渐近收敛界,所有常数均以维度和逆温度的形式明确跟踪,改进了当前基于次梯度的朗之万算法的已知速率。我们还为相关优化问题提供了超额风险估计。我们针对GPT-2系列大型语言模型(LLM)的正则化预训练势验证了带有明确常数的假设,而SG-TULA的提升逐坐标变体对前者的预训练效果可与微调后的AdamW和Muon相媲美,目前尚无针对后者的可比非渐近保证。
英文摘要
We study the problem of sampling from target distributions whose potentials are simultaneously non-smooth, subject to superlinear gradient growth, and non-convex. We introduce the Subgradient Tamed Unadjusted Langevin Algorithm (SG-TULA), a discretisation of the Langevin diffusion that operates directly on subgradients, without relying on computationally demanding smoothing procedures. To handle the superlinear regime, taming techniques are employed to produce a stable, explicit scheme. We derive non-asymptotic convergence bounds in Wasserstein-2 distance, with all constants tracked explicitly in terms of dimension and inverse temperature, improving upon the currently known rates for subgradient-based Langevin algorithms. We further provide excess risk estimates for the associated optimisation problem. We verify the assumptions, with explicit constants, for the regularized pretraining potential of a LLM in the GPT-2 lineage and the boosted coordinate-wise variant of SG-TULA pretrains the former competitively against finetuned AdamW and Muon, for which no comparable non-asymptotic guarantees are presently available.
Comments53 pages