因果Transformer在上下文长度上的普适性与泛化性
Universality and Generalization of Causal Transformers Across Context Lengths
- Doshisha University(同志社大学)
- RIKEN AIP(理化学研究所人工智能研究中心)
- Rice University(莱斯大学)
- CNRS(法国国家科学研究中心)
- ENS(巴黎高等师范学院)
- PSL Université(巴黎文理研究大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文研究单个因果Transformer能否在不同上下文长度下统一近似token映射,提出基于Hölder连续性的框架,给出泛化误差界,并通过物理时间序列实验验证。
AI中文摘要:
长上下文是现代Transformer系统的核心,但大多数表达性结果针对每个固定序列长度选择不同的网络。我们研究一个掩码Transformer能否在任意长度序列上统一近似因果的token到token映射,这些序列采样自一个固定的归一化时间范围。为了关联不同采样分辨率,我们用$\alpha$-Hölder序列或更一般的连续模来建模token。我们跨分辨率的连续性概念刻画了因果族,这些族允许通过具有长度无关参数的单一Transformer在这些紧致输入类上实现统一近似。该结果扩展到无限长度的平均场极限,其中token形成连续曲线,掩码注意力变为因果时间积分。对于有界回归,目标映射满足使用正则测试函数定义的$\beta$-光滑稳定性条件,定量近似产生一个泛化界:在适当大小的有界权重Transformer上进行精确经验风险最小化,从$N$个独立同分布的标记序列中得到的均方根预测误差为$O((\log\log N/\log N)^{\beta/(d+2)})$。该界在同一采样分布上以固定置信度成立,其中$d$为token维度,且没有最大长度因子。最后,在物理时间序列上的实验支持了在观测尺度上的Hölder正则token模型,具有依赖于数据集的拟合指数,而文本输入嵌入则提供了一个对比案例。原生和密集采样、打乱控制和细化检查界定了这一经验正则性区域。
英文摘要:
Long contexts are central to modern transformer systems, but most expressivity results choose a different network for each fixed sequence length. We study whether one masked transformer can approximate causal token-to-token maps uniformly over sequences of arbitrary length sampling a fixed normalized horizon. To relate sampling resolutions, we model tokens by $α$-Hölder sequences or, more generally, a common modulus of continuity. Our notion of continuity across resolutions characterizes the causal families admitting uniform approximation on these compact input classes by a single transformer with length-independent parameters. The result extends to the infinite-length mean-field limit, where tokens form continuous curves and masked attention becomes a causal time integral. For bounded regression with target maps satisfying a $β$-smooth stability condition defined using regular test functions, quantitative approximation yields a generalization bound: exact empirical risk minimization over suitably sized bounded-weight transformers gives root mean-square prediction error $O((\log\log N/\log N)^{β/(d+2)})$ from $N$ iid labeled sequences. The bound holds at fixed confidence on the same sampling distribution, with $d$ the token dimension and no maximum-length factor. Finally, experiments on physical time series support the Hölder-regular token model at observed scales, with dataset-dependent fitted exponents, whereas text input embeddings provide a contrasting case. Native and dense sampling, shuffled controls, and refinement checks delimit this empirical regularity regime.