发表机构
National University of Singapore; Princeton University; Nanyang Technological University(新加坡国立大学; 普林斯顿大学; 南洋理工大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文研究梯度预训练如何学习决策桩阈值,提出二参数软注意力模型,证明其误差界并揭示参数发散机制,同时指出单头模型的局限性。
AI 中文摘要
估计决策阈值需要定位未知边界附近的观测值。我们研究了基于梯度的预训练如何在具有固定特征和不等式方向的二参数软注意力模型中学习这一统计规则。预训练使用带标签的上下文及其真实阈值;新阈值必须仅从上下文中推断。在大分辨率初始化下,对包含 $m$ 个任务、每个任务 $n$ 个样本的数据集执行常数步长梯度下降,会产生一个冻结估计器,对于每个固定的内部阈值和每个新上下文大小 $N$,其误差为 $\widetilde O((m\wedge n)^{-1}+N^{-1})$。这两项分别对应有限预训练精度和新上下文定位精度。其机制是协调的参数发散:群体训练校准相对标签和特征分数,然后以 $t^{1/4}$ 的速率增加注意力尺度,使群体阈值误差达到 $O(t^{-1/4})$。为了将该机制转移到固定的有限语料库,我们控制了相对于连续参数尺度上进展缩减方向的梯度误差。这验证了增长训练区间的可行性,而无需长期跟踪群体轨迹。我们还识别了单头模型的边界局限性,并从统计学角度解释了反射对称化可能实现的效果。
英文摘要
Estimating a decision threshold requires locating observations near an unknown boundary. We study how gradient-based pretraining learns this statistical rule in a two-parameter softmax-attention model with a fixed feature and inequality direction. Pretraining uses labeled contexts and their true thresholds; a fresh threshold must be inferred from context alone. Under a large-resolution initialization, constant-step gradient descent on $m$ tasks with $n$ examples each produces a frozen estimator with error $\widetilde O((m\wedge n)^{-1}+N^{-1})$ for each fixed interior threshold and every fresh-context size $N$. The two terms separate finite-pretraining accuracy from fresh-context localization. The mechanism is coordinated parameter divergence: population training calibrates the relative label and feature scores, then increases the attention scale as $t^{1/4}$, giving population threshold error $O(t^{-1/4})$. To transfer this mechanism to a fixed finite corpus, we control gradient errors relative to the shrinking directions of progress at successive parameter scales. This certifies a growing training interval without requiring long-time tracking of the population trajectory. We also identify the boundary limitation of the one-head model and explain statistically what a reflected symmetrization could achieve.