AI 中文总结
该研究针对Deepseek Engram模块与分词器耦合的问题,通过替换哈希方式构建联合嵌入空间,实现了分词器无关性,同时保持了相当的性能。
AI 中文摘要
Deepseek的Engram是一种条件记忆模块,用于在大型语言模型中权衡存储与推理的关系。不过该模块依赖分词器级的N元语法哈希来进行Engram嵌入查找,导致其与所用分词器紧密耦合:采用不同分词器的模型必须从头训练自身的Engram嵌入。为提升Engram嵌入的可复用性,我们对哈希例程进行了修改,使采用不同分词器的Engram模型能够兼容。我们不再对不相交的N元语法空间进行建模,而是将N元语法视为从所有分词器可能的字节序列中采样潜在有用字节序列的方法;我们将基于异或(XOR)的哈希替换为通用多项式哈希,并构建了跨N元语法的联合嵌入空间。本研究探讨了可能的权衡关系,结果表明这一简单替换可产生相当的性能,且实现了与分词器无关的特性:即字节等价的分词序列具有哈希等价性。
英文摘要
Deepseek's Engram, a conditional memory module, was introduced to trade-off storage versus reasoning in large language models. However, the module relies on token-level $N$-gram hashing for Engram embedding lookup, introducing a tight coupling to the tokenizer used: a model with a different tokenizer would have to train its own Engram embeddings from scratch. To improve the reusability of Engram embeddings, we propose a change to the hashing routine, enabling compatibility between Engram models using different tokenizers. Instead of modelling disjoint $N$-gram spaces, we treat $N$-gram as a method to sample potentially useful byte sequences, from all possible byte sequences across tokens. We replace the XOR-based hashing with the general polynomial hashing with a joint embedding space across $N$. This work investigates the possible trade-offs and shows that this simple substitution produces comparable performance and achieves tokenizer-agnosticism: hash equivalence for byte-equivalent token sequences.
CommentsPreprint, 7 pages