发表机构
Beijing University of Posts and Telecommunications; Kuaishou Technology(北京邮电大学; 快手科技)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对Transformer记忆容量受限问题,提出Lngram v2解耦记忆与骨干网络参数,提升可扩展性,在30B参数VLMs上性能提升且降低参数,离散ID可稳定关联语义。
AI 中文摘要
Transformer 缺乏原生查找机制,需要重复进行密集计算以识别和复用局部静态模式。Lngram v1 通过离散潜在N元语法寻址引入了与分词器无关的条件记忆,但其记忆容量与骨干网络宽度耦合,因参数和激活成本高而限制了可扩展性。我们提出的 Lngram v2 将路由数量、记忆维度和骨干网络宽度解耦,并引入上下文感知分组查询注意力读出,以独立扩展记忆容量。零值 Sink 和反事实代理梯度进一步在保留硬离散寻址的同时,提升了读出选择性和路由可训练性。针对不同规模的视觉-语言模型(VLMs)开展的实验显示出一致的性能提升,包括成功扩展至300亿参数模型。与 Lngram v1 相比,Lngram v2 在维持或提升语言建模性能的同时,大幅降低了总记忆参数和激活记忆参数。进一步分析表明,其离散ID保留了连续隐藏状态的大量语义结构,仅通过ID即可实现语义恢复,且在不同数据集间具备稳定的ID-语义关联。这些结果确立了 Lngram v2 作为一种高效且可扩展的潜在条件记忆机制,其离散寻址还为分析模型内部表示提供了结构化接口。
英文摘要
Transformers lack a native lookup mechanism, requiring repeated dense computation to recognize and reuse local static patterns. Lngram v1 introduces tokenizer-independent conditional memory through discrete latent n-gram addressing, but its memory capacity is coupled with the backbone width, limiting scalability due to high parameter and activation costs. We propose Lngram v2, which decouples the number of routes, memory dimension, and backbone width, and introduces a context-aware grouped-query attention readout to scale memory capacity independently. A zero-value Sink and counterfactual surrogate gradients further improve readout selectivity and routing trainability while preserving hard discrete addressing. Experiments across vision--language models (VLMs) of different scales show consistent improvements, including successful scaling to a 30B-parameter model. Compared with Lngram v1, Lngram v2 substantially reduces both total and activated memory parameters while maintaining or improving language modeling performance. Further analysis shows that its discrete IDs preserve substantial semantic structure of continuous hidden states, enabling semantic recovery from IDs alone and stable ID--semantic associations across datasets. These results establish Lngram v2 as an efficient and scalable latent conditional memory mechanism whose discrete addresses also provide a structured interface for analyzing internal model representations.