arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.21000cs.AI

Naju:一种具有独立保留和写入功能的用于长序列记忆的原生离散状态空间模型

Naju: A Native Discrete State-Space Model with Independent Retention and Writing for Long-Sequence Memory

  • Korea Institute of Energy Technology (KENTECH)(韩国能源技术研究所)

机构由 AI 辅助整理,请以论文原文为准。

Hyuk Lim, Seunghyun Yoon

AI总结:

研究针对长序列记忆跟踪中保留与覆盖的矛盾需求,提出Naju模型,将循环更新分解为离散极点、写入增益及映射。该模型在训练长度4倍时仍兼具保留与覆盖能力,在多任务中结合长距记忆与良好性能,优于Mamba基线。

AI中文摘要:

长序列记忆跟踪对循环状态提出了两个相互矛盾的要求:在长时间内近乎无损地保留存储的绑定,以及主动覆盖陈旧的绑定。在我们的诊断套件中,最强的有效基线往往只能很好地解决其中一个方面。像Mamba这样的连续时间参数化状态空间模型(SSMs)通过对连续时间系统进行零阶保持离散化来获得其离散递归;我们认为这种迂回对于记忆跟踪来说是不必要的,直接对离散转移进行参数化。Naju(原生自适应连接单元)将循环更新分解为一个显式离散极点(一个学习到的遗忘门$f_n$)、一个独立的写入增益$i_n$以及与输入相关的写入/读取映射。由于Sigmoid极点满足$0<f_n<1$,每个冻结的局部坐标通过构造是舒尔稳定的,并且在均匀有界假设下,完整的时变递归满足渐消记忆/BIBO界,无需稳定性正则化。我们形式化了耦合设计的关键结构限制:任何非扩张互补单门递归通过$|r|+w\leq1$将有效保留$r$和写入增益$w$联系起来,因此近乎完全保留会迫使写入较弱;将$f_n$与$i_n$解耦消除了这个约束。从经验上看,Naju是唯一在训练长度为4倍时在保留和覆盖方面都保持强大的评估模型。在诊断套件之外,我们在WikiText - 103语言建模、长距离竞技场和多查询关联召回上评估了Naju。在这些设置中,Naju始终将强大的长距离记忆与有竞争力或卓越的性能相结合,在主要比较中优于Mamba基线,同时与Transformer保持竞争力并保留线性时间、线性内存缩放。

英文摘要:

Long-sequence memory tracking places two opposing demands on a recurrent state: near-lossless retention of stored bindings over long horizons, and active overwriting of stale ones. In our diagnostic suite, the strongest efficient baselines tend to solve only one side well. Continuous-time-parameterized state-space models (SSMs) such as Mamba obtain their discrete recurrence by zero-order-hold discretization of a continuous-time system; we argue that this detour is unnecessary for memory tracking and parameterize the discrete transition directly. Naju (Native Adaptive Junction Unit) factorizes the recurrent update, schematically $x_n = f_n\odot x_{n-1} + i_n\odot(B_n u_n)$, into an explicit discrete pole (a learned forget gate $f_n$), an independent write gain $i_n$, and input-dependent write/read maps. Since the sigmoid pole satisfies $0<f_n<1$, each frozen local coordinate is Schur-stable by construction, and the full time-varying recurrence satisfies a fading-memory/BIBO bound under uniform boundedness assumptions, with no stability regularizer. We formalize the key structural limitation of coupled designs: any non-expansive complementary single-gate recurrence ties the effective retention $r$ and write gain $w$ through $|r|+w\le 1$, so near-complete retention forces weak writing; decoupling $f_n$ from $i_n$ removes this constraint. Empirically, Naju is the only evaluated model that remains strong on both retention and overwriting at 4x the training length. Beyond the diagnostic suite, we evaluate Naju on WikiText-103 language modeling, Long Range Arena, and multi-query associative recall. Across these settings, Naju consistently combines strong long-range memory with competitive or superior performance, outperforming the Mamba baselines in the principal comparisons while remaining competitive with the Transformer and preserving linear-time, linear-memory scaling.

↑