发表机构
UCLA; UT Austin; Doshisha University; RIKEN AIP(加州大学洛杉矶分校; 德克萨斯大学奥斯汀分校; 同志社大学; 理化学研究所人工智能研究中心)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文从数学上研究变换器定义的测度到测度算子,证明其将亚高斯输入映射为亚高斯输出且关于1-Wasserstein距离Hölder连续,并给出误差传播估计、交叉注意力的正则性与样本复杂度分析及逼近保证,为亚高斯数据上的变换器奠定稳定性与有限样本理论基础。
AI 中文摘要
变换器在各种领域表现出令人印象深刻的经验成功,但其理论基础仍不够完善。这项工作是对变换器定义的测度到测度算子的数学研究。我们证明变换器将亚高斯输入映射到亚高斯输出;这确保了任意长度的softmax算子组合是良定义的。然后我们证明变换器在适当的亚高斯输入空间上关于1-Wasserstein距离是Hölder连续的。这使我们能够建立沿变换器在亚高斯输入与其经验近似之间的误差传播估计。我们还研究了交叉注意力机制的均值场类比,它是一个从一对概率测度到单一概率测度的算子。我们表明交叉注意力在其两个输入参数中表现出不同的Hölder正则性和样本复杂度。最后,我们将我们的结果应用于推导测度到测度变换器的逼近保证。总之,这些结果为亚高斯数据上的变换器提供了坚实的稳定性和有限样本理论。
英文摘要
Transformers have exhibited impressive empirical success across various domains, but their theoretical foundations remain less developed. This work constitutes a mathematical study of the measure-to-measure operators defined by transformers. We show that transformers map sub-Gaussian inputs to sub-Gaussian outputs; this ensures that taking arbitrary-length compositions of the softmax operator is well-defined. We then show that transformers are Hölder continuous with respect to the 1-Wasserstein distance on appropriate spaces of sub-Gaussian inputs. This allows us to establish estimates on the error propagation along a transformer between a sub-Gaussian input and its empirical approximation. We also study a mean-field analog of the cross-attention mechanism, which is an operator from a pair of probability measures to a single probability measure. We show that cross-attention exhibits different Hölder regularity and sample-complexity in its two input arguments. Last, we apply our results to deduce approximation guarantees for measure-to-measure transformers. Together, these results provide a firm stability and finite-sample theory for transformers on sub-Gaussian data.