arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Transformer训练与Wasserstein分布鲁棒性的最优控制泛化界

Generalization Bounds on Optimal Control for Transformer Training and Wasserstein Distributional Robustness

Kağan Akman, Naci Saldi, Serdar Yüksel

arXiv 2607.27975首次发表:更新:

发表机构

Bilkent University(比尔肯特大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究推导了Transformer训练的有限样本泛化界,通过将训练问题转化为马尔可夫控制问题,结合量化模型与Wasserstein分布鲁棒优化,建立了Transformer泛化与分布鲁棒控制的关联。

AI 中文摘要

我们推导了采用动态规划递归训练的Transformer的有限样本泛化界。基于Transformer动力学的双提升测度值公式,我们将数据集视为经验输入-输出测度对的概率律,从而将训练问题解释为有限时间马尔可夫控制问题。接着,我们分析了通过量化状态、动作和测度状态空间得到的量化模型,利用有限度量空间上经验律的集中不等式以及值函数的Lipschitz稳定性估计推导显式有限样本泛化界。这些界以显式近似误差为代价被传递到基础模型。最后,我们表明相同的机制可得到训练问题的分布鲁棒控制公式,将Transformer泛化与Wasserstein分布鲁棒优化联系起来。

英文摘要

We derive finite-sample generalization bounds for Transformers trained with dynamic programming recursions. Building on the doubly lifted, measure-valued formulation of Transformer dynamics, we view data sets as probability laws on pairs of empirical input-output measures, allowing us to interpret the training problem as a finite-horizon Markovian control problem. We then analyze a quantized model, derived by quantizing the state, action, and measure-state spaces, and derive explicit finite-sample generalization bounds using concentration inequalities for empirical laws on finite metric spaces together with a Lipschitz stability estimate for the value function. These bounds are transferred to the base model at the cost of an explicit approximation error. Finally, we show that the same machinery yields a distributionally robust control formulation of the training problem, connecting Transformer generalization to Wasserstein distributionally robust optimization.

Comments25 pages

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑