发表机构
The Pennsylvania State University(宾夕法尼亚州立大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究如何实现无参数自适应稀疏注意力,核心方法是利用gzip压缩率识别内容块来动态生成稀疏掩码,主要贡献是在语言建模实验中性能超越其他模型,无参数添加且随序列长度优势扩大、收敛更快。
AI 中文摘要
数据自适应稀疏注意力掩码在长序列上表现出色。现有自适应方法通常需要额外参数等。本文表明经典数据压缩能提供无额外参数的有效掩码信号。通过计算每块gzip压缩率识别非冗余内容块并选择性地进行远程注意力。在PG - 19字节级语言建模实验中,该方法无参数添加且性能优于其他模型,随着序列长度增加优势扩大且收敛更快。
英文摘要
Data-adaptive sparse attention masks substantially outperform fixed patterns (e.g., BigBird and Longformer) and can even exceed dense attention on long sequences. Existing adaptive approaches---including SBM-Transformer, Dynamic Mask Attention, and NSA---typically require additional learnable parameters, custom gradient estimators, or specialized CUDA kernels. We show that classical data compression provides an effective masking signal with \textbf{no additional parameters}. By computing per-block gzip compression ratios, we identify non-redundant content blocks and route long-range attention selectively through them. Intuitively, blocks that gzip cannot compress contain information not predictable from local repetition, making them natural long-range attention targets. Because the compression profile is input-dependent, the resulting sparse mask adapts dynamically to content without learned parameters, auxiliary losses, or custom kernels. On PG-19 byte-level language modeling at 92M parameters with 8K context, our method achieves 1.71 bits-per-byte (BPB), outperforming dense attention (2.89), BigBird (2.34), Longformer (3.21), and a reimplemented SBM-Transformer (3.38)---the only learned-mask baseline---by up to 1.67 BPB while adding no parameters. The advantage grows with sequence length, with the gap over BigBird widening from 0.05 BPB at 4K context to 0.63 BPB at 8K, while convergence is 3.3$\times$ faster.