AI 中文总结
通过将注意力提升到辛相空间,绕过SHT单层归纳障碍,实现精确辛且可部署的注意力,并在91.3M参数下观察到归纳涌现。\n
AI 中文摘要
我们在后RoPE基底上的线性、单步、因果、双线性、辛一致设计类中,绕过了Sanford-Hsu-Telgarsky(SHT)单层归纳障碍,方法是将注意力提升到辛相空间,这镜像了Hairer对Stormer-Verlet的提升。该提升脱离了SHT计数论证的前提,而非其界限本身。重新表述将障碍展现为滤波器阶数差距:单层双线性得分实现了一个联合阶数为(0,0)的z变换,而归纳判别器需要键侧阶数≥1。将辛上剪切M_γ:(q,p)↦(q+γp,p)应用于后RoPE查询和键流可闭合此差距。我们证明此提升在因子化子类中是唯一的,在算子级别精确辛,并且需要后RoPE放置;在显式的仅T_4高斯约化中,我们推导出一个闭式双分支归纳相变,在r=0.9876处保持,且零拟合参数。该定律是可解析求解的极限,而非稳健预测:恢复T_3通道在d_k=64时将γ_c从1.030移至0.569并消除交叉。可部署性通过精确推导得出:KV缓存不变,前缀复用和推测解码得以保留,每层每token开销为6d FLOPs,INT8余量最多增加log_2(1+2γ)比特,融合内核未修改,且不添加参数。在91.3M参数下,超临界扫描定位了一个涌现带:在γ=0.80时,归纳在平均717步内形成3/3种子,而在γ=0时为2/3种子和2700步。不利结果同样直接报告:仅键半提升达到0.949,而对称算子为0.811,因此若归纳精度为目标,半提升是更好的构造。四十一个笔记本和结果文件作为辅助材料随附。
英文摘要
We circumvent the Sanford-Hsu-Telgarsky (SHT) single-layer induction obstruction within a linear, one-step, causal, bilinear, symplectically consistent design class on the post-RoPE substrate, by lifting attention onto a symplectic phase space, mirroring Hairer's lift of Stormer-Verlet. The lift exits the premise of the SHT counting argument rather than the bound itself. The reframing exhibits the obstruction as a filter-order gap: a one-layer bilinear score realises a $z$-transform of joint order $(0,0)$, whereas the induction discriminator requires key-side order $\geq 1$. Applying the symplectic upper shear $M_γ:(q,p)\mapsto(q+γp,p)$ to the post-RoPE query and key streams closes it. We prove this lift is unique within the factorised subclass, exactly symplectic at operator level, and requires post-RoPE placement; and in an explicit $T_4$-only Gaussian reduction we derive a closed-form two-branch induction phase transition, held out at $r=0.9876$ with zero fitted parameters. That law is an analytically solvable limit, not a robust prediction: restoring the $T_3$ channel moves $γ_c$ at $d_k=64$ from 1.030 to 0.569 and removes the crossover. Deployability follows by exact derivation: the KV cache is unchanged, prefix reuse and speculative decoding are preserved, overhead is $6d$ FLOPs per token per layer, INT8 headroom grows by at most $\log_2(1+2γ)$ bits, fused kernels are unmodified, and no parameters are added. At 91.3M parameters a supercritical sweep locates an emergence band: induction forms 3/3 seeds at $γ=0.80$ in a mean of 717 steps, against 2/3 seeds and 2700 steps at $γ=0$. Adverse results are reported as directly: a key-only half-lift reaches 0.949 against 0.811 for the symmetric operator, so if induction accuracy is the objective, the half-lift is the better construction. Forty-one notebooks and result files ship as ancillary material.