arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.15126cs.CLcs.AI

MoME:用于上下文感知稀疏查找的混合记忆嵌入

MoME: Mixture-of-Memory Embeddings for Context-Aware Sparse Lookup

  • University of British Columbia(不列颠哥伦比亚大学)
  • Vector Institute for AI(向量人工智能研究所)

机构由 AI 辅助整理,请以论文原文为准。

Muchen Li, Leonid Sigal, Renjie Liao

AI总结:

本文提出MoME,一种上下文感知的混合记忆嵌入机制,通过门控选择多槽位以区分多义令牌,在多个主干网络上优于基线,并展现出语义可解释性。

AI中文摘要:

高效扩展大型语言模型推动了稀疏容量机制的发展,如混合专家模型,以及最近的条件下记忆:即通过令牌索引的嵌入表,以廉价的参数化查找来增强主干网络。现有的记忆嵌入方法通过表面形式的确定性函数进行检索,这将同一令牌的不同上下文语义(例如,python 作为编程语言与作为动物)压缩到单个固定条目中。我们引入了混合记忆嵌入(MoME),一种上下文感知的记忆机制,将每个令牌的单一记忆行替换为 M 个槽位的混合,并使用基于隐藏状态的学习门控在每个位置选择要读取的槽位。在 nanochat、Llama-3/MobileLLM 和 Qwen3 主干网络上的受控预训练实验中,MoME 在等参数和等训练 FLOPs 设置下优于 Value Embedding、Bigram 和 STEM 基线,在亚十亿规模下显示出更有前景的记忆大小扩展趋势,并在训练和推理中保持高效。对多义令牌的定性路由分析进一步表明,学习到的混合体展现出一定程度的语义可解释性,将相同表面令牌在不同语义下分配到不同的记忆槽位。

英文摘要:

Scaling large language models efficiently has motivated sparse capacity mechanisms such as Mixture-of-Experts and, more recently, conditional memory: token-indexed embedding tables that augment the backbone with cheap parametric lookups. Existing memory-embedding methods retrieve via a deterministic function of the surface form, which collapses different contextual senses of the same token (e.g., python the language vs. the animal) into a single fixed entry. We introduce Mixture of Memory Embeddings (MoME), a context-aware memory mechanism that replaces each token's single memory row with a mixture of M slots and uses a learned gate over the hidden state to choose which slots to read at each position. In controlled pretraining experiments across nanochat, Llama-3/MobileLLM, and Qwen3 backbones, MoME improves over Value Embedding, Bigram, and STEM baselines in iso-parameter and iso-training-FLOP settings, shows a more promising memory-size scaling trend at sub-billion scale, and remains efficient in training and inference. Qualitative routing analyses on polysemous tokens further suggest that the learned mixture exhibits a degree of semantic interpretability, dispatching the same surface token to distinct memory slots under different senses.

补充信息

↑