DeepSeek-V4.1-Flash:突破KV缓存压缩极限
DeepSeek-V4.1-Flash: Pushing the Limits of KV Cache Compression
浏览论文内容
中文总结 AI 辅助
DeepSeek-V4.1-Flash通过CED架构、CSA2跨层KV缓存复用与FP4缓存,将全局KV缓存降至每token 890字节,显著降低部署成本,同时性能优于基线。
中文摘要 AI 辅助
长时程智能体的广泛采用使得模型工作负载日益偏向输入密集型。尽管先前的工作已大幅降低了长上下文计算成本,但预填充阶段的计算开销仍然高昂,且庞大的KV缓存持续对HBM和SSD的容量及数据传输带宽造成压力。这些计算、存储和带宽需求共同构成了进一步降低部署成本的主要瓶颈。为解决这一挑战,我们推出了DeepSeek-V4.1-Flash,一个多模态混合专家(MoE)模型,拥有552B主干参数,支持高达一百万token的上下文。凭借其因果编码器-解码器(CED)架构,该模型在解码阶段每token激活16B参数,但在预填充阶段仅激活8B参数,显著提升了智能体工作负载的成本效率。为突破KV缓存压缩的极限,DeepSeek-V4.1-Flash将压缩稀疏注意力2(CSA2)中的跨层KV缓存复用与FP4 KV缓存相结合。这些设计将其全局KV缓存占用(始终驻留于HBM)降至每token 890字节,约为DeepSeek-V4-Flash相应占用量的1/4。此外,通过一项名为SWA有界重放的专用部署优化,DeepSeek-V4.1-Flash将其持久KV缓存占用(始终驻留于SSD或主机内存)降至约为DeepSeek-V4-Flash的1/8。尽管其KV缓存占用大幅减小,该模型的性能仍显著优于基线。此外,我们精简了DeepSeek-V4架构,并引入了若干高效的架构扩展。我们在包含45T token的多模态语料库上预训练了DeepSeek-V4.1-Flash,并进行了全面的后训练,在多种基于文本和多模态的智能体场景中均取得了强劲性能。模型检查点可在该https URL获取。
英文摘要
The widespread adoption of long-horizon agents has made model workloads increasingly input-heavy. Although prior work has substantially reduced the cost of long-context computation, prefill remains computationally expensive, and large KV caches continue to strain HBM and SSD capacity and data-transfer bandwidth. Together, these compute, storage, and bandwidth demands constitute the primary bottleneck to further lowering deployment costs. To address this challenge, we introduce DeepSeek-V4.1-Flash, a multimodal Mixture-of-Experts (MoE) model with 552B backbone parameters and support for contexts of up to one million tokens. With its Causal Encoder-Decoder (CED) architecture, the model activates 16B parameters per token during decode but only 8B parameters during prefill, substantially improving cost efficiency for agentic workloads. To push the limits of KV cache compression, DeepSeek-V4.1-Flash combines cross-layer KV cache reuse in Compressed Sparse Attention 2 (CSA2) with FP4 KV caching. These designs reduce its global KV cache footprint (always in HBM) to 890 bytes per token, roughly 1/4 of the corresponding footprint of DeepSeek-V4-Flash. Further, through a dedicated deployment optimization known as SWA Bounded Replay, DeepSeek-V4.1-Flash reduces its persistent KV cache footprint (always on SSD or in host memory) to roughly 1/8 of that of DeepSeek-V4-Flash. Despite its much smaller KV cache footprint, the model delivers substantially better performance than the baseline. In addition, we streamline the DeepSeek-V4 architecture and introduce several efficient architectural extensions. We pretrain DeepSeek-V4.1-Flash on a multimodal corpus comprising 45T tokens and conduct comprehensive post-training, yielding strong performance across diverse text-based and multimodal agentic scenarios. Model checkpoints are available at https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash.
发表机构
- DeepSeek-AI(深度求索人工智能)
机构由 AI 辅助整理,请以论文原文为准。