arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

用于LLM推理中可扩展KV缓存管理的光子-CXL存储设备

A Photonic-CXL Memory Appliance for Scalable KV Cache Management in LLM Inference

Jing Ding, Yash Nishant, Chandrish Ambati, Jyothsna Kamati, Trung Diep

arXiv 2607.27187首次发表:更新:

AI 中文总结

该研究针对LLM推理的存储墙问题,提出Marvell光子-CXL混合存储设备,实现32 TB共享内存,延迟降超50%,多轮对话首token时间提升6.6倍,解决了电气CXL池的部署痛点。

AI 中文摘要

大规模LLM推理面临存储墙问题,KV缓存需求达数十TB,且需数百GB每秒的带宽,但当前无存储层级可同时满足这两项需求。对多代GPU系统及各类LLaMA模型的表征测试显示,主机内存检索的速度是重计算的100倍,但仅支持数十个并发长上下文用户。电气CXL池在理论上可弥合这一差距,但存在交换机延迟、线缆传输距离限制及功率扩展问题,阻碍了TB级部署的实际落地。本文提出Marvell Photonic Fabric(PF)Memory Appliance,这是一种光子-CXL混合架构,用无源光纤交叉连接替代电气交换机,通过无交换机的全交叉拓扑在16台主机间提供32 TB共享内存。仿真结果表明,与电气CXL池相比,其延迟降低超过50%;针对多轮对话工作负载,PF Memory Appliance将首token生成时间提升6.6倍,消除了缓存驱逐的性能断崖。

英文摘要

LLM inference at scale faces a memory wall. The KV cache demands tens of terabytes at hundreds of gigabytes per second, yet no current memory tier delivers both at once. Characterization across multi-generation GPU systems with various LLaMA models shows host memory retrieval achieves up to 100x speedup over re-computation but supports only tens of concurrent long-context users. Electrical CXL pooling theoretically bridges this gap, but switch latency, cable reach limits, and power-scaling issues prevent practical TB-scale deployments. We present the Marvell Photonic Fabric Memory Appliance, a photonic-CXL hybrid architecture replacing electrical switches with a passive fiber shuffle to deliver 32 TB shared memory across 16 hosts via a switch-free full- crossbar topology. Emulation results demonstrate over 50 percent latency reduction versus electrical CXL pools. Simulation results show that the PF Memory Appliance eliminates cache eviction cliffs by improving time-to-first-token by 6.6x for multi-turn conversations workloads.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑