边缘语言模型的结构化记忆:通过O(1) SSM状态注入实现持久上下文与语料库检索
Structured Memory for Edge Language Models: Persistent Context and Corpus Retrieval via O(1) SSM State Injection
浏览论文内容
中文总结 AI 辅助
该研究针对边缘语言模型提出PRECOG与SMC机制,将SSM预填充成本压缩至O(1),在1.2B参数的TENNs-LLM上实现约4500倍预填充加速,达到与RAG相当的答案质量。
中文摘要 AI 辅助
检索增强生成(RAG)会产生与检索上下文长度成正比的预填充成本,且对于Transformer骨干网络而言,键值缓存(KV-cache)会随每个生成的令牌增长。状态空间模型(SSM)从结构上避免了第二种成本;我们消除了第一种成本,将每个查询的预填充从O(L_context)压缩至O(1)。我们提出PRECOG(预计算上下文注入),这是一种利用SSM独有特性的检索机制:固定大小、与位置无关的循环隐藏状态是模型所读取所有内容的完整摘要。PRECOG在离线阶段将文档语料预编码为SSM隐藏状态,并在查询时直接注入匹配度最高的状态,完全绕过上下文重新摄入过程。相同的状态注入机制还支持SMC(结构化记忆整合):一种具有认知领域聚类、可调整的保真度-存储权衡以及O(1)会话初始化的分层持久记忆,它将短期情景状态整合为长期语义记忆,并在查询时将两者与检索到的语料库状态融合。我们在TENNs-LLM上验证了该系统,这是一个具有192 KB隐藏状态的12亿参数门控SSM语言模型。PRECOG达到了上下文RAG的答案质量,在边缘硬件上将预填充延迟从约27秒降低至小于6毫秒,实现了约4500倍的加速,跨越了从不可用到可交互的阈值。该机制对于Transformer KV缓存而言在架构上是不可能的,因为后者与位置纠缠且随上下文长度线性增长。
英文摘要
Retrieval-augmented generation (RAG) imposes a prefill cost proportional to retrieved context length, and -- with Transformer backbones -- a KV-cache that grows with each generated token. State-Space Models (SSMs) avoid the second cost by construction; we eliminate the first, collapsing prefill from $O(L_{context})$ to $O(1)$ per query. We introduce PRECOG (Pre-Computed Context Injection), a retrieval mechanism that exploits a property unique to SSMs: the fixed-size, position-agnostic recurrent hidden state is a complete summary of everything the model has read. PRECOG pre-encodes document corpora offline as SSM hidden states and injects the best-matching state directly at query time, bypassing in-context re-ingestion entirely. The same state-injection mechanism enables SMC (Structured Memory Consolidation): a hierarchical persistent memory with cognitive-domain clustering, an adjustable fidelity-vs-storage dial, and $O(1)$ session initialization, which consolidates short-term episodic states into long-term semantic memory and fuses both with retrieved corpus states at query time. We demonstrate the system on TENNs-LLM, a 1.2B-parameter gated-SSM language model with a 192 KB hidden state. PRECOG matches in-context RAG answer quality, reducing prefill latency from $\sim$27 s to $<$6 ms on edge hardware -- a $\sim$4500$\times$ speedup that crosses the threshold from unusable to interactive. The mechanism is architecturally impossible for Transformer KV-caches, which are position-entangled and grow linearly with context length.
发表机构
- BrainChip Inc.(BrainChip公司)
机构由 AI 辅助整理,请以论文原文为准。