arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

RedKnot-MLA:面向DeepSeek-V4长上下文服务的多头离线-在线复用

RedKnot-MLA: Multi-Head Offline-Online Reuse for DeepSeek-V4 Long-Context Serving

Yang Liu, Zhaokai Luo, Huayi Jin, Ruozhou He, Chenchen Hong, Mingxiao Ma, Biao Zhang, Zhiyong Wang, Boyu Wang, Guanjie Chen, Yifei Liu, Tao Xie, Junhao Hu

arXiv 2609.07008首次发表:更新:

发表机构

Huawei Cloud; Shanghai Jiao Tong University; Xiaohongshu Inc.; Peking University; Beijing Tongming Lake Information Technology Application Innovation Center(华为云; 上海交通大学; 小红书公司; 北京大学; 北京通明湖信息技术应用创新中心)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出RedKnot-MLA系统,通过离线-在线复用实现DeepSeek-V4长上下文服务,利用MLA头感知分解与位置修复,在保持精度的同时显著加速TTFT并节省算术计算。

AI 中文摘要

多头潜在注意力(MLA)通过一个打包的潜在KV流暴露了许多逻辑查询头。这种表示在内存上高效,但它消除了传统逐头复用所假设的物理每头缓存边界。我们提出了我们的系统,即RedKnot头感知复用原理的DeepSeek-V4实现。每个不可变文档在规范位置零处离线处理;认证的局部头贡献被保留为MLA-Off。在服务时,查询侧RoPE重定位恢复文档的请求位置,一个小的全局头集和受保护的局部令牌行被重新计算为MLA-Online,两条路径在单个共享输出投影之前合并。打包的MLA潜在表示从不被拆分。DeepSeek-V4-Flash使用37个可复用层和56/8的局部/全局划分,给出75.29%的分析逻辑头行上限;Pro-0813配置文件使用55层和112/16个头,给出78.89%。冻结的Flash操作点显示热工件TTFT加速2.02-3.84倍。在256K下,存档的三数据集研究报告聚合F1变化+3.24个百分点,EM变化+4.16个百分点,以及78.7-79.5%的分析主要算子算术节省,而一个数据集下降2.81个F1点。一个单独的作者报告的256K热工件QPS测量约为2.0倍;由于其原始并发轨迹未包含在此包中,我们将其标记为初步而非存档证据。我们描述了分解、位置修复、令牌行闭包、稀疏MoE支持、TP8集成以及解释这些结果所需的测量边界。

英文摘要

Multi-head latent attention (MLA) exposes many logical query heads through one packed latent KV stream. This representation is memory efficient, but it removes the physical per-head cache boundary assumed by conventional head-wise reuse. We present our system, a DeepSeek-V4 realization of RedKnot's head-aware reuse principle. Each immutable document is processed offline at canonical position zero; certified Local-head contributions are retained as MLA-Off. At serving time, query-side RoPE relocation restores the document's request position, a small Global-head set and protected Local token rows are recomputed as MLA-Online, and the two paths are merged before a single shared output projection. The packed MLA latent is never split. DeepSeek-V4-Flash uses 37 reusable layers and a 56/8 Local/Global partition, giving a 75.29% analytic logical head-row ceiling; the Pro-0813 profile uses 55 layers and 112/16 heads, giving 78.89%. Frozen Flash operating points show hot-artifact TTFT speedups of 2.02-3.84x. At 256K, the archived three-dataset study reports an aggregate F1 change of +3.24 percentage points, an EM change of +4.16 points, and a 78.7-79.5% analytic major-operator arithmetic saving, while one dataset decreases by 2.81 F1 points. A separate author-reported 256K hot-artifact QPS measurement is approximately 2.0x; because its raw concurrency trace is not included in this bundle, we mark it as preliminary rather than archived evidence. We describe the factorization, position repair, token-row closure, sparse-MoE support, TP8 integration, and the measurement boundaries needed to interpret these results.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑