AI 中文总结
该研究针对LLM推理的存储墙问题,提出Marvell光子-CXL混合存储设备,实现32 TB共享内存,延迟降超50%,多轮对话首token时间提升6.6倍,解决了电气CXL池的部署痛点。
AI 中文摘要
大规模LLM推理面临存储墙问题,KV缓存需求达数十TB,且需数百GB每秒的带宽,但当前无存储层级可同时满足这两项需求。对多代GPU系统及各类LLaMA模型的表征测试显示,主机内存检索的速度是重计算的100倍,但仅支持数十个并发长上下文用户。电气CXL池在理论上可弥合这一差距,但存在交换机延迟、线缆传输距离限制及功率扩展问题,阻碍了TB级部署的实际落地。本文提出Marvell Photonic Fabric(PF)Memory Appliance,这是一种光子-CXL混合架构,用无源光纤交叉连接替代电气交换机,通过无交换机的全交叉拓扑在16台主机间提供32 TB共享内存。仿真结果表明,与电气CXL池相比,其延迟降低超过50%;针对多轮对话工作负载,PF Memory Appliance将首token生成时间提升6.6倍,消除了缓存驱逐的性能断崖。
英文摘要
LLM inference at scale faces a memory wall. The KV cache demands tens of terabytes at hundreds of gigabytes per second, yet no current memory tier delivers both at once. Characterization across multi-generation GPU systems with various LLaMA models shows host memory retrieval achieves up to 100x speedup over re-computation but supports only tens of concurrent long-context users. Electrical CXL pooling theoretically bridges this gap, but switch latency, cable reach limits, and power-scaling issues prevent practical TB-scale deployments. We present the Marvell Photonic Fabric Memory Appliance, a photonic-CXL hybrid architecture replacing electrical switches with a passive fiber shuffle to deliver 32 TB shared memory across 16 hosts via a switch-free full- crossbar topology. Emulation results demonstrate over 50 percent latency reduction versus electrical CXL pools. Simulation results show that the PF Memory Appliance eliminates cache eviction cliffs by improving time-to-first-token by 6.6x for multi-turn conversations workloads.