基于数组数学(MoA)的结构化解码注意力DNF推导、KV缓存累积、GQA/MQA和OpenACC内核
MoA-Structured Decode Attention DNF Derivation, KV-Cache Accumulation, GQA/MQA, and OpenACC Kernel
浏览论文内容
中文总结 AI 辅助
该研究利用数组数学为Transformer注意力推导内存最优推理工件,包括单查询解码DNF、GPU内核、多步KV缓存以及GQA和MQA,通过特定方法实现内存优化,且程序经与PyTorch验证。
中文摘要 AI 辅助
我们使用数组数学(MoA)为Transformer注意力推导了四个内存最优推理工件。每个工件都直接来自于前向传递的表示范式(DNF),查询行索引固定为当前解码步骤。工件包括:一个单查询解码DNF,通过ψ归约代数消除K^T缓冲区;一个C/OpenACC图形处理单元(GPU)内核,具有操作范式(ONF)步长算法和硬件合并内存访问;一个多步KV缓存;以及通过ψ选择推导的分组查询注意力(GQA)和多查询注意力(MQA)。所有程序都与PyTorch的scaled_dot_product_attention进行了验证。
英文摘要
We derive four memory-optimal inference artifacts for transformer attention using the Mathematics of Arrays (MoA), each following directly from the forward-pass Denotational Normal Form (DNF) of with the query-row index fixed to the current decode step. The artifacts are: (1)~a single-query decode DNF in which the $ψ$-reduction eliminates the $K^\top$ buffer algebraically, achieving $(d_k + nd_k+ nd_v+ d_v)\times4\,{B}$ Dynamic Random Access Memory (DRAM) traffic result numerically verified to $\|{err}\|_\leq2\times10^{-7}$; (2)~a C/OpenACC Graphics Processing Unit (GPU) kernel with Operational Normal Form (ONF) stride arithmetic and hardware-coalesced memory access, verified to $\|\mathrm{err}\|_\infty=0$ (exact IEEE-754 floating-point arithmetic); (3)~a multi-step KV-cache with $O(d_k+d_v)$ per-step append via MoA concatenation $\#$; and (4)~Grouped-Query Attention (GQA) and Multi-Query Attention (MQA) derived via $ψ$-selection, achieving a proven $\frac {h_q} { h_{kv} }$ reduction in KV traffic. All programs are verified against PyTorch scaled_dot_product_attention.
发表机构
- University at Albany, SUNY(纽约州立大学奥尔巴尼分校)
- LACL, Université Paris-Est Créteil(巴黎东部克雷泰伊大学LACL)
机构由 AI 辅助整理,请以论文原文为准。