Q-First:在块的最小改动下实现注意力与前馈的并发
Q-First: Most of Attention Needs Only the Query in Disaggregated LLM Decoding
浏览论文内容
中文总结 AI 辅助
该研究提出Q-First协议,通过交换LLM解码器块的注意力与前馈子层顺序消除依赖,实现并发,验证其精度并训练8种变体证明改动影响有界且符合要求。
中文摘要 AI 辅助
拆分式大语言模型(LLM)服务将KV缓存扫描部署在内存优化硬件上,投影层与前馈网络部署在计算优化硬件上,由此继承了解码器块中双方都不愿承担的依赖关系:注意力计算先执行,前馈网络需消耗其输出,因此在同个序列内,一方工作时另一方会空闲。常规的修复方式是为每个额外的飞行序列保留一份驻留KV缓存,而这正是拆分设备的初衷。我们转而消除这种依赖关系:扫描仅需要查询(query),交换两个子层的顺序可使查询在计算侧仍有工作时就可用,当前的键(key)和值(value)作为缓存写入,无需等待。我们将得到的解码过程表述为协议,证明其可在标准内核上运行,并在训练好的检查点上进行端到端验证,相对误差达3.2×10^-3——无需新算子、改变形状或新硬件。随后,我们仅调整注意力的读取位置,其余参数固定,以2个随机种子各训练8种块变体。每字节比特数(bits per byte)在计算最优值的3%时,用于衡量改动对训练的干扰程度而非最终性能,因此我们关注数值大小而非排名。在5个前馈网络不消耗自身注意力的块中,任意读取点与未改动块的差异均不超过0.0026 bits per byte,小于同一模型在另一随机种子下的自身差异(0.0066),而相同设置下,子层交换的差异达前者的25倍。提前移动查询是测量无法检测到的改动,这正是协议所需的。改动的影响有界:从网络输入投影每一层查询会增加+0.0974,在两个随机种子下均反驳了预先注册的阈值,因此查询可提前一个前馈网络读取,无法再提前。
英文摘要
Disaggregated LLM serving puts the KV-cache sweep on memory-optimised hardware and the projections and feed-forward on compute-optimised hardware, then inherits from the decoder block a dependency neither device wants: attention runs first and the feed-forward consumes its output, so within one sequence each side idles while the other works. The usual repair costs one resident KV cache per extra sequence in flight, which is what motivated separating the devices at all. We remove the dependency instead. The sweep needs only the query, and exchanging the two sub-layers makes that query available while the compute side still has work to do, so the two run concurrently; the current key and value follow as a cache write nothing waits on. We state the decode as a protocol, show that it runs on stock kernels, and verify it end to end on a trained checkpoint to a relative error of 3.2x10^-3 -- with no new operator, no changed shape and no new hardware. We then train the block 8 ways at two seeds each, varying only where the attention reads and holding everything else fixed. At three per cent of compute-optimal a lead in bits per byte measures how much a change disturbed training rather than what it reaches, so we read magnitudes and not rankings. Among the 5 blocks whose feed-forward does not consume their own attention, no read point differs from the one that moves nothing by more than 0.0026 bits per byte -- smaller than the gap between an arm and itself at a second seed, 0.0066 -- while the same runs resolve a sub-layer exchange 25 times as large. Moving the query early is a change the measurement cannot find, which is what the protocol needs. The reach is bounded: projecting every layer's query from the network's input costs +0.0974, refuting a pre-registered threshold at both seeds, so a query may be read one feed-forward early and no further back.