发表机构
Georgia Institute of Technology; Emory University(佐治亚理工学院; 埃默里大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究为Transformer算子学习建立了理论框架,推导了逼近误差界和泛化幂律缩放定律,揭示了内在低维结构对收敛速率的影响,并经数值实验验证。
AI 中文摘要
Transformer已成为学习物理系统解算子的强大架构。经验上,当数据规模和模型规模增大时,预测误差会减小,这表明存在神经缩放行为。然而,对于基于Transformer的算子学习,这种缩放定律的理论理解仍然有限。在这项工作中,我们开发了一个理论框架来刻画基于Transformer的算子学习的逼近误差和泛化误差。我们的分析建立在一种局部到全局的逼近原理之上,该原理与softmax注意力机制自然对齐,并产生离散化不变的输出函数。在逼近理论方面,我们推导了基于Transformer的算子学习对于Hölder正则算子的通用逼近误差。在泛化理论方面,我们建立了预测误差与训练数据规模之间的幂律缩放关系。由缩放指数表示的收敛速率明确反映了输入和输出域的维度、底层函数和算子的正则性,以及关键地,输入函数类的内在维度。通过利用这种内在低维结构,我们的分析得出了算子学习的幂律泛化速率,这与现有使用前馈神经网络的算子学习分析中出现的对数型幂律速率形成对比。数值实验验证了预测的幂律缩放,并确认收敛速率随输入函数类的内在维度系统性变化。
英文摘要
Transformers have emerged as powerful architectures for learning solution operators of physical systems. Empirically the prediction error has been observed to decrease when the data size and model size increase, suggesting neural scaling behavior. Yet a theoretical understanding of such scaling laws for transformer-based operator learning remains limited. In this work, we develop a theoretical framework for characterizing the approximation and generalization errors of transformer-based operator learning. Our analysis builds on a local-to-global approximation principle that is naturally aligned with the softmax attention mechanism and yields discretization-invariant output functions. On approximation theory, we derive a universal approximation error of transformer-based operator learning for Hölder-regular operators. On generalization theory, we establish a power scaling law between the prediction error and the training data size. The rate of convergence represented by the scaling exponent explicitly reflects the dimensions of the input and output domains, the regularity of the underlying functions and operators, and crucially, the intrinsic dimension of the input function class. By exploiting this intrinsic low-dimensional structure, our analysis yields a power-law generalization rate for operator learning, in contrast to the logarithmic-type power-law rates appearing in existing analyses of operator learning with feedforward neural networks. Numerical experiments validate the predicted power-law scaling and confirm that the convergence rate varies systematically with the intrinsic dimension of the input function class.
Comments47 pages, 5 figures