BF1:一种用于高效长上下文Transformer的因果二元稀疏注意力改造方案
BF1: A Causal Dyadic Sparse-Attention Retrofit for Efficient Long-Context Transformers
浏览论文内容
中文总结 AI 辅助
本文提出BF1,一种结合局部邻域、全局首块与对数间隔历史块的稀疏注意力方案,改造Qwen3-0.6B部分注意力层后,在长上下文场景下显著降低模型延迟,且语言建模性能优于基线方法。
中文摘要 AI 辅助
即使采用高度优化的精确内核,密集因果注意力在长上下文场景下仍计算成本高昂。本文研究了BF1,这是一种确定性的块对齐二元稀疏注意力路由,结合了小型精确局部邻域、全局首块以及对数间隔的历史块。该路由与现有对数稀疏和膨胀注意力模式相关;本文的贡献包括正确性门控预训练模型改造、匹配的拓扑控制研究,以及将每层稀疏度与整体模型延迟关联的系统表征。对于固定块宽度,每个转换层使用O(n log n)个选定的令牌交互,具有O(log n)的图通信深度。在NVIDIA RTX PRO 6000 Blackwell GPU上,优化的BF16实现方案在2K至4K令牌范围内优于密集注意力,在32K令牌时每层预填充速度提升达10.91倍。将28层Qwen3-0.6B注意力层中的8层改造后,在8K、16K、32K令牌场景下,模型预热的首次令牌生成时间分别降低7.7%、11.3%、15.3%,其余密集层使完整模型仍保持渐近二次复杂度。在匹配的1000步、1638.4万令牌的适配协议下,BF1在3个训练种子中排名第一:平均报告困惑度为1.68639,而匹配的静态随机非局部图为1.69154,密集继续训练为1.69258,等预算局部滑动为1.81505。在种子1234下,打包报告的配对区间显示,密集继续训练(Dense-CT)比BF1高0.3169%-0.4055%,静态随机图比BF1高0.2441%-0.3642%。这些结果证实BF1是一种可复现的稀疏算子和选择性改造原语,具有实际长上下文系统价值。本文评估了数值正确性、选定交互的缩放性、内核性能、部分模型推理以及匹配的下一个令牌语言建模。
英文摘要
Dense causal attention remains expensive at long context even when implemented with highly optimized exact kernels. We study BF1, a deterministic block-aligned dyadic sparse-attention route that combines a small exact local neighborhood, a global first block, and logarithmically spaced historical blocks. The route is related to prior log-sparse and dilated attention patterns; our contribution is a correctness-gated pretrained-model retrofit, a matched topology-control study, and a systems characterization that connects per-layer sparsity to whole-model latency. For fixed block width, every converted layer uses O(n log n) selected token interactions and has O(log n) graph communication depth. On an NVIDIA RTX PRO 6000 Blackwell GPU, an optimized BF16 implementation crosses dense attention between 2K and 4K tokens and reaches a 10.91x per-layer prefill speedup at 32K. Retrofitting eight of 28 Qwen3-0.6B attention layers lowers warm whole-model time to first token by 7.7%, 11.3%, and 15.3% at 8K, 16K, and 32K, respectively, while the remaining dense layers keep the complete model asymptotically quadratic. Under a matched 1,000-step, 16.384M-token adaptation protocol, BF1 ranks first across three training seeds: mean report perplexity is 1.68639 versus 1.69154 for a matched static-random nonlocal graph, 1.69258 for dense continued training, and 1.81505 for equal-budget local sliding. At seed 1234, the packed-report paired interval places Dense-CT 0.3169-0.4055% above BF1 and static-random graph 17 0.2441-0.3642% above BF1. These results establish BF1 as a reproducible sparse operator and selective retrofit primitive with real long-context systems value. This paper evaluates numerical correctness, selected-interaction scaling, kernel performance, partial-model inference, and matched next-token language modeling.