MoNe:用于高效长上下文推理的模块化神经记忆
MoNe: Modular Neural Memory for Efficient Long Context Inference
浏览论文内容
中文总结 AI 辅助
研究提出MoNe,一种可附加到冻结预训练Transformer的轻量级模块化神经记忆,通过两阶段设计实现高效长上下文推理,在128K令牌场景下较ICL降低约80%计算与内存开销,泛化性能优于ICL。
中文摘要 AI 辅助
我们提出了MoNe,这是一种轻量级模块化神经记忆,可附加到任何冻结的预训练Transformer上,无需重新训练即可实现长上下文推理。MoNe通过对具有层本地化梯度更新的快速权重神经记忆网络进行测试时学习,以固定大小的片段读取上下文;在推理时,该记忆仅从查询令牌生成键和值,无需重新读取上下文令牌。这种两阶段设计将推理成本与上下文长度解耦,实现了O(N)的预处理成本和O(1)的查询成本,且峰值GPU内存不会随N增长。在128K令牌的情况下,与上下文学习(ICL)相比,MoNe将计算量和峰值GPU内存均降低了约80%,仅产生6.4%的参数开销。MoNe可泛化到远超主干原生窗口的上下文长度,在RULER的“干草堆中的针”和单词提取基准上表现出色,而ICL则在此类任务中性能急剧下降。
英文摘要
We present MoNe, a lightweight modular neural memory that attaches to any frozen pretrained Transformer to enable long-context inference without retraining. MoNe reads context in fixed-size segments via test-time learning of fast-weight neural memory networks with layer-localized gradient updates; at inference, the memory generates keys and values from the query tokens alone, with no context tokens re-read. This two-phase design decouples inference cost from context length, achieving $O(N)$ preprocessing and $O(1)$ query cost with peak GPU memory that does not grow with $N$. At 128K tokens, MoNe reduces both compute and peak GPU memory by approximately 80% compared to ICL with only 6.4% parameter overhead. MoNe generalizes to context lengths far beyond the backbone's native window, achieving strong performance on needle-in-a-haystack and word extraction benchmarks from RULER, where ICL degrades sharply.
发表机构
- Qualcomm AI Research(高通人工智能研究部门)
- Qualcomm Technologies, Inc(高通技术公司)
机构由 AI 辅助整理,请以论文原文为准。