From Buffers to Registers: Unlocking Fine-Grained FlashAttention with Hybrid-Bonded 3D NPU Co-Design
从缓冲区到寄存器:通过混合键合3D NPU协同设计解锁细粒度FlashAttention
机构 * SKLP, Institute of Computing Technology, Chinese Academy of Sciences, Beijing, China(科学技术信息研究所、中国科学院、北京) ; University of Chinese Academy of Sciences, Beijing, China(中国科学院大学、北京)
AI总结 本文提出3D-Flow架构和3D-FlashAttention方法,通过混合键合3D堆叠空间加速器实现细粒度FlashAttention,降低能耗并提升速度。
Comments Accepted to DATE 2026