arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.15627cs.DCcs.ARcs.PF

DeepSeek-V4-Flash 在 AMD gfx90a 上的实现:正确性恢复与推理性能工程

DeepSeek-V4-Flash on AMD gfx90a: Correctness Recovery and Inference Performance Engineering

  • HKUST(GZ)(香港科技大学(广州))

机构由 AI 辅助整理,请以论文原文为准。

Siming Huang

AI总结:

本文在 AMD MI250 (gfx90a) 上实现 DeepSeek-V4-Flash 推理,修复路由专家 W2 布局错误,通过 FP4/FP8 优化和内核调优,实现约 74.5 tok/s 解码和 2,234 tok/s 预填充。

AI中文摘要:

我们展示了 DeepSeek-V4-Flash 在采用 gfx90a/CDNA2 架构的 AMD Instinct MI250 GPU 上的启用、正确性恢复和性能工程。该系统在 SGLang 中集成了原生 safetensors 加载、张量并行和专家并行、FP4 路由混合专家计算、FP8 稠密投影、稀疏注意力、HIP 图执行以及兼容 OpenAI 的服务。最初快速执行路径被发现因路由专家 W2 布局不匹配而在数值上不正确。我们识别了输出置换,在加载时修复权重布局,并在进一步优化前建立固定令牌和基于哈希的正确性检查。在修正后的路径上,通过打包 FP4 权重、INT8 激活量化、CDNA2 点积指令、对等读取全归约和拓扑感知内核几何形状,解码性能得到提升。预填充通过 CDNA2 MFMA 内核、改进的打包权重重用、降低稀疏注意力开销、更大的块和重新调整的专家排序来加速。在四个 MI250 GCD 上,TP4/EP1 原生自回归解码达到约 74.5 tok/s,而 4,604 令牌提示的 TTFT 为 2.061-2.062 秒,即约 2,234 输入 tok/s。结果表明,CDNA2 上高效的 DeepSeek-V4-Flash 推理不仅受限于内存带宽,还受限于 FP4 执行格式不匹配、低 MFMA 利用率和逐层同步成本。

英文摘要:

We present the enablement, correctness recovery, and performance engineering of DeepSeek-V4-Flash inference on AMD Instinct MI250 GPUs using the gfx90a/CDNA2 architecture. The system integrates native safetensors loading, tensor and expert parallelism, FP4 routed mixture-of-experts computation, FP8 dense projections, sparse attention, HIP graph execution, and OpenAI-compatible serving within SGLang. An initially fast execution path was found to be numerically incorrect because of a routed-expert W2 layout mismatch. We identify the output permutation, repair the weight layout at load time, and establish fixed-token and hash-based correctness checks before further optimization. On the corrected path, decode performance is improved through packed FP4 weights, INT8 activation quantization, CDNA2 dot-product instructions, peer-read all-reduce, and topology-aware kernel geometry. Prefill is accelerated using CDNA2 MFMA kernels, improved packed-weight reuse, reduced sparse-attention overhead, larger chunks, and retuned expert sorting. On four MI250 GCDs, TP4/EP1 native autoregressive decode reaches approximately 74.5 tok/s, while a 4,604-token prompt reaches 2.061-2.062 s TTFT, or approximately 2,234 input tok/s. The results show that efficient DeepSeek-V4-Flash inference on CDNA2 is limited not only by memory bandwidth, but also by FP4 execution-format mismatch, low-M utilization, and per-layer synchronization costs.

↑