arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.18966cs.LGcs.NE

底层自动机:加法输入通路是Householder线性RNN中状态追踪的寄生吸引子

The Automaton Underneath: The Additive Input Pathway Is a Parasitic Attractor for State Tracking in Householder Linear RNN

Gunner Levi Howe

首次发表
浏览论文内容

中文总结 AI 辅助

本研究通过预注册因果消融发现,Householder线性RNN中的加法输入通路是寄生吸引子,移除它可使模型在长度泛化上达到精确自动机学习,并揭示表示定律与任务格式的隐藏作用。

中文摘要 AI 辅助

具有输入相关Householder乘积转移(DeltaNet/DeltaProduct类)的线性RNN可以证明地表示困难的状态追踪自动机,然而训练后的模型无法进行长度泛化——最近的工作将此差距归因于优化问题,但没有因果解释。我们在一项预注册的、架构内的因果消融研究中给出了这样的解释:同一个模型删除一项——加法输入注入 $b_t = W_b e_t$。有 $b_t$ 时,模型在奇偶校验、$S_4$、$A_5$ 和不可解 $S_5$ 字问题上拟合长度32但在分布外崩溃(在 $S_5$ 的位置512处准确率为0.20)。没有它——输入仅通过正交转移作用——同一架构学习到精确的自动机:在训练长度的16倍处,中位数准确率为1.00,在我们陈述并测试的表示定律所允许的每个宽度上:每个token的Householder因子最小数量等于任务生成器在格式固定表示中的最大反射长度(奇偶校验1,$S_4$ 3,$A_5$ 和 $S_5$ 4);低于此值,什么都无法拟合。与DeltaProduct在 $n_h{=}2$ 时的 $S_4$/$A_5$(群元素分类,$SO(3)$ 实现)的对比表明该定律是表示相对的:任务格式是状态追踪基准中的隐藏变量。两个预注册分支定位了机制。(i) 在 $W_b{=}0$ 的已验证精确解处初始化,Adam增长加法路径并将模型拉离精确解;$-b$ 控制保持在1.00。(ii) 我们注册的预测,即拟合通过 $b_t$ 路由,触发了其终止标准:所有49个拟合种子在 $W_b{:=}0$ 下保留域内拟合——并且在定律最小宽度下,推理时清零 $W_b$ 恢复精确泛化(奇偶校验5/5,$S_5$ 5/5,$S_4$ 4/5,$A_5$ 4/5)。加法通路是寄生的:它先破坏,然后隐藏一个正确学习的自动机。所有202次运行均预注册;所有数字均可从工件重新生成。

英文摘要

Linear RNNs with input-dependent Householder-product transitions (DeltaNet/DeltaProduct-class) can provably represent hard state-tracking automata, yet trained models fail to length-generalize -- a gap recent work attributes to optimization, without a causal account. We give one, in a pre-registered, within-architecture causal ablation: the same model with one term deleted -- the additive input injection $b_t = W_b e_t$. With $b_t$, models fit length 32 and collapse out-of-distribution on parity, $S_4$, $A_5$, and non-solvable $S_5$ word problems (0.20 at position 512 on $S_5$). Without it -- input acting only through the orthogonal transitions -- the same architecture learns the exact automaton: median accuracy 1.00 at 16x the training length, at every width admitted by a representation law we state and test: the minimal number of Householder factors per token equals the maximal reflection length of the task's generators in the format-pinned representation (parity 1, $S_4$ 3, $A_5$ and $S_5$ 4); below it, nothing fits. The contrast with DeltaProduct's $S_4$/$A_5$ at $n_h{=}2$ (group-element classification, $SO(3)$ realization) shows the law is representation-relative: task format is a hidden variable in state-tracking benchmarks. Two pre-registered arms locate the mechanism. (i) Initialized at a verified-exact solution with $W_b{=}0$, Adam grows the additive path and pulls the model off the exact solution; $-b$ controls stay at 1.00. (ii) Our registered prediction that the fit routes through $b_t$ fired its kill criterion: all 49 fitting seeds retain in-domain fit under $W_b{:=}0$ -- and at law-minimal width, zeroing $W_b$ at inference restores exact generalization (parity 5/5, $S_5$ 5/5, $S_4$ 4/5, $A_5$ 4/5). The additive pathway is parasitic: it destabilizes, then conceals, a correctly learned automaton. All 202 runs pre-registered; all numbers regenerate from artifacts.

↑