AI 中文总结
该研究通过全栈表征发现,将高带宽闪存(HBF)直接替代SSD作为瞬态KV卸载层会降低LLM服务性能,HBF需作为感知复用的协调层使用而非SSD替代品。
AI 中文摘要
高带宽闪存(High-Bandwidth Flash, HBF)通过在封装本地的宽接口后堆叠NAND,提供了闪存级容量,且读取延迟和带宽远优于固态硬盘(SSD)。这使得人们很想保留SSD风格的Mooncake KV卸载栈,仅将后端层替换为HBF。我们使用扩展的TokenSim、四个完整的两小时通义千问-百炼(Qwen-Bailian)生产轨迹、五个密集型和混合专家模型,以及H100/B200配置文件测试了这种替换。结果显示,服务性能反而变差而非变好,成本效益模型解释了原因。更快的远端层仅在读取I/O是服务瓶颈、读取量超过写入量且交付带宽可持续时才有用,这三个条件必须同时满足,而瞬态KV在每种情况下都不满足。封装交换会占用GPU近层的容量和带宽,因此在H100和B200上,平均端到端延迟上升2--5.5倍,最大服务水平目标(SLO)有效吞吐量下降1.1--2.7倍。服务对HBF自身的读写延迟几乎不敏感,且基底管芯近内存计算不会提高闪存层在关键路径中的占比。两层结构使近层保持复用,并为HBF提供了写入密集型流,因此在每条轨迹中写入量都超过读取量。3D-ICE模型显示,该流会使栈在远低于峰值带宽时达到热极限,且TLC层比容量匹配的SSD池磨损更快。更快的设备会导致系统更慢,因为封装放弃的资源多于介质带来的收益。HBF本身不是问题,将其作为更快的SSD用于瞬态KV才是问题所在。它应作为选择性、感知复用、写入预算受限且热协调的层用于服务,而非作为SSD的直接替代品。
英文摘要
A faster storage device should make serving faster. We find the opposite. High-Bandwidth Flash (HBF) stacks NAND behind a wide, package-local interface, promising flash-scale capacity with far lower read latency and higher bandwidth than an SSD. The obvious move is to keep an SSD-style Mooncake KV-offloading stack and swap in HBF underneath. We built that system and measured it: an extended TokenSim, four complete two-hour Qwen-Bailian production traces, five dense and mixture-of-experts models, and H100/B200 profiles. The upgrade backfires. Average end-to-end latency rises 2--5.5$\times$ and maximum SLO goodput falls 1.1--2.7$\times$ across H100 and B200, so the faster device yields a slower system. A cost-benefit model explains the paradox: a faster far tier pays off only when read I/O is the bottleneck, reads outweigh writes, and delivered bandwidth is sustainable. Transient KV violates all three at once. Buying flash through the package costs GPU near-tier capacity and bandwidth, while HBF's own read/write latency barely matters: scaling it 3.75$\times$ moves latency less than 1\%. Worse, the two-tier hierarchy keeps reuse in the near tier and hands HBF a relentless write-heavy stream. Writes outnumber reads on every trace, so a 3D-ICE model shows the stack hits its thermal limit well below peak bandwidth, and a TLC tier wears out sooner than the SSD pool it replaced. The device is fine; the drop-in deployment is not. HBF sucks as an SSD replacement for transient KV, but earns its place in LLM serving when used selectively with reuse-aware placement, write budgeting, and thermal coordination.