AI 中文总结
研究固定状态序列模型与注意力模型间的中间地带,提出基于DP均值聚类规则的可学习稀疏缓存,有静态和惊喜自适应两种形式,在基准测试中表现良好,能端到端学习,在多真实流中验证了不同项目属性。
AI 中文摘要
固定状态序列模型将无界的过去压缩到有界状态中,这限制了它们的关联记忆,大致在状态维度;注意力通过为每个令牌保留键值条目来突破这一限制,但计算量为二次方,且缓存随序列增长。我们研究了中间地带:一种稀疏缓存,仅在输入新颖时分配一个插槽,因此其大小跟踪不同项目的数量而非令牌数量。分配规则是DP均值聚类规则,即狄利克雷过程混合的小方差极限,但不是用作潜在变量推理,而是用作深度循环主干的键值记忆运算符。我们开发了两种形式,一种是具有固定浓度的静态缓存,另一种是惊喜自适应变体,其浓度跟随最近的新颖率。在具有冗余的受控关联记忆基准测试中,我们表明该缓存仅存储不同项目时就能匹配全注意力记忆,在召回率与大小前沿上优于固定预算逐出缓存,并且在状态空间主干上,它在任何测试模型中以最低内存回答召回查询和远程聚合。分配是端到端可学习的:仅在任务损失上训练的双参数新颖阈值门能准确恢复规则,而过参数化的门则失败,所以起作用的因素是归纳偏差而非容量。证据是一系列适度规模的受控机制研究,在四个真实流(推荐、系统日志、临床事件和保险索赔)上证实了不同项目属性;在一项配套研究中进行了真实主干、真实语料库的语言验证。
英文摘要
Fixed-state sequence models compress an unbounded past into a bounded state, which caps their associative recall at roughly the state dimension; attention escapes the cap by keeping a key-value entry for every token, at quadratic compute and a cache that grows with the sequence. We study the middle ground: a sparse cache that allocates a slot only when an input is novel, so its size tracks the number of distinct items rather than the number of tokens. The allocation rule is the DP-means clustering rule, the small-variance limit of a Dirichlet-process mixture, used not as latent-variable inference but as the key-value memory operator for a deep recurrent backbone. We develop it in two forms, a static cache with a fixed concentration and a surprise-adaptive variant whose concentration follows the recent novelty rate. On a controlled associative-recall benchmark with redundancy we show that the cache matches full-attention recall while storing only the distinct items, that it dominates a fixed-budget eviction cache on the recall-versus-size frontier, and that on a state-space backbone it answers both a recall query and a long-range aggregate at the lowest memory of any model tested. The allocation is learnable end to end: a two-parameter novelty-threshold gate trained on the task loss alone recovers the rule exactly, whereas an over-parameterized gate fails, so the operative ingredient is the inductive bias rather than capacity. The evidence is a family of controlled mechanism studies at modest scale, with the distinct-items property confirmed on four real streams (recommendation, systems logs, clinical events, and insurance claims); a real-backbone, real-corpus language validation is pursued in a companion study.
Comments16 pages, 9 figures. Companion paper on event-log applications forthcoming