AI 中文总结
针对有限精度下并行扫描导致的扰动,提出神经有限状态机(NFSM),通过非仿射扫描兼容RNN实现长度无关的有限状态跟踪,在代数与文本任务上超越仿射基线。
AI 中文摘要
学习鲁棒且可扩展的有限状态跟踪是序列处理的基础。虽然线性循环神经网络(RNN)、线性注意力和状态空间模型通过仿射递推实现了可扩展的并行训练,但它们的理论表达能力保证假设了理想化算术,并不适用于有限精度,在有限精度下,使它们变快的并行扫描本身就是扰动的一个来源。我们在有限精度下形式化有限状态跟踪,并刻画长度无关性:在精度成本不随长度增长的情况下,跟踪在每个序列长度上都保持正确。我们表明,长度无关的状态跟踪需要在单个映射内存在两种竞争动态:收缩以抑制数值扰动,以及分离以保持不同状态彼此区分。我们证明了仿射递推(每一步只提供一个速率来同时扮演这两个角色)在有限精度下最多能实现确定自动机。我们不把扫描兼容性视为对更新映射的限制,而是将其重新解释为一种计算预算,并引入神经有限状态机(NFSM):一种非仿射、扫描兼容的RNN,专为长度无关的有限状态跟踪而构建。在涵盖阿贝尔群和非阿贝尔群、不可逆幺半群以及文本状态跟踪任务的合成基准上,仿射基线在每个非确定任务上都失败,其中大多数在几百步内失败。单个NFSM层反而能学习每个代数任务的精确转换表,这证明了在测试长度之外的正确性,而堆叠的NFSM在文本任务上每个测试长度都保持完美准确率。
英文摘要
Learning robust and scalable finite-state tracking is fundamental to sequence processing. While linear recurrent neural networks (RNNs), linear attention, and state space models enable scalable parallel training through affine recurrences, their theoretical expressivity guarantees assume idealized arithmetic and do not extend to finite precision, where the parallel scan that makes them fast is itself a source of perturbation. We formalize finite-state tracking at finite precision and characterize length independence: tracking that stays correct at every sequence length, at a precision cost that does not grow with the length. We show that length-independent state tracking requires two competing dynamics within a single map: contraction to suppress numerical perturbations and separation to keep distinct states apart. We prove that affine recurrences, which offer a single rate at each step to serve both roles, realize at most definite automata at finite precision. Instead of treating scan compatibility as a restriction on the update map, we reinterpret it as a computational budget and introduce the Neural Finite-State Machine (NFSM): a nonaffine, scan-compatible RNN built for length-independent finite-state tracking. On synthetic benchmarks spanning abelian and nonabelian groups, noninvertible monoids, and textual state-tracking tasks, affine baselines fail on every nondefinite task, most of them within a few hundred steps. A single NFSM layer instead learns the exact transition tables of every algebraic task, which certifies correctness beyond the tested lengths, and a stack of NFSMs keeps perfect accuracy on the textual tasks at every tested length.
Commentshttps://julienbrandoit.github.io/length-independent-state-tracking/