BCMT:分块因果记忆Transformer
BCMT: Blockwise Causal Memory Transformer
浏览论文内容
中文总结 AI 辅助
该研究提出BCMT架构,通过分块局部自注意力与指数因果记忆解耦长程依赖建模,在1024token上下文的语言建模中,性能与密集Transformer相当,且训练吞吐量更高、内存消耗更低。
中文摘要 AI 辅助
Transformer架构依赖密集自注意力对长程依赖进行建模,但该机制的复杂度随序列长度呈二次增长。我们提出BCMT(Blockwise Causal Memory Transformer,分块因果记忆Transformer),这是一种用于长上下文语言建模的架构,它将局部token交互与全局上下文传播解耦。在局部块内独立应用密集因果自注意力,同时每个块通过指数因果记忆生成聚合的自适应摘要。该记忆随后被注入回token表示中,无需依赖显式全局注意力即可实现长程上下文信息的高效传播。与标准Transformer和循环记忆架构不同,BCMT既不维持远距离token间的密集交互,也不使用学习到的记忆状态。其记忆机制完全可并行化,且与密集自注意力的标准实现兼容。在上下文长度达1024个token的语言建模实验中,BCMT的验证性能与密集Transformer相当,同时显著提升了训练吞吐量并降低了内存消耗。消融研究进一步证实,这些改进源于所提出的记忆机制。这些结果表明,由分块摘要构建的指数因果记忆为长上下文语言建模提供了密集全局注意力机制的有效替代方案。
英文摘要
Transformer architectures rely on dense self-attention to model long-range dependencies, but this mechanism exhibits quadratic complexity with respect to sequence length. We introduce BCMT (Blockwise Causal Memory Transformer), an architecture for long-context language modeling that decouples local token interactions from global context propagation. Dense causal self-attention is applied independently within local blocks, while each block produces an adaptive summary aggregated through an exponential causal memory. This memory is subsequently injected back into the token representations, enabling efficient propagation of long-range contextual information without relying on explicit global attention. Unlike standard Transformers and recurrent memory architectures, BCMT maintains neither dense interactions between distant tokens nor learned memory states. Its memory mechanism is fully parallelizable and remains compatible with standard implementations of dense self-attention. Experiments on language modeling with context lengths of up to 1024 tokens show that BCMT achieves validation performance comparable to that of Dense Transformers while significantly improving training throughput and reducing memory consumption. An ablation study further confirms that these improvements arise from the proposed memory mechanism. These results demonstrate that an exponential causal memory constructed from block summaries provides an effective alternative to dense global attention mechanisms for long-context language modeling.