发表机构
MIT; Vizuara(麻省理工学院; 维祖拉公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究对比LLM服务中张量并行与KV压缩的成本性能,发现压缩在多数场景成本更低,仅当模型规模超80GB卡对应36B参数时需用张量并行,且压缩提升容量、张量并行降低延迟。
AI 中文摘要
当大语言模型(LLM)服务部署耗尽KV缓存空间时,有两种成熟的解决方案:张量并行将权重和KV缓存分片到2、4或8个设备,通过每层的全归约操作换取内存余量,但硬件成本随设备数量增加而上升;算法领域则通过原位压缩缓存,采用KV量化和淘汰机制,在单GPU上运行,以小幅性能损失为代价。压缩相关论文报告内存比率,并行扩展相关论文报告吞吐量曲线,几乎无人将两者置于同一成本维度。我们将张量并行配置(并行度1至8)和KV压缩配置(16/8/4位,保留率低至0.25)置于同一成本归一化维度,使用基于A100、A40和H100硬件校准的仿真器,以每百万令牌的成本对延迟作图,寻找成本等价的交叉点,但未找到该点。在两个模型(Llama-2 7B和70B)、三种GPU类型及所有可构建的内存缓解级别下,压缩方案的成本低1.20至2.00倍。80GB设备上的7B模型无法在自身上下文窗口内耗尽KV预算,两种策略的决策边界是模型规模与设备内存的比值,约为80GB卡对应36B参数。低于该阈值时,压缩方案占优,额外GPU基本是浪费支出;高于该阈值时,张量并行不再是选择,而是准入门槛:Llama-2-70B在单个A100上无论何种KV设置均不可行,因为绑定资源是权重,而KV压缩不涉及权重。张量并行是唯一能降低延迟的手段(压缩因批处理竞争使每令牌延迟升高8%至93%),而压缩是唯一能提升每美元容量的手段(提升16.5倍,而8倍GPU支出仅提升1.21倍)。
英文摘要
When an LLM serving deployment runs out of KVcache room, there are two well-established ways out. Tensor parallelism shards the weights and the KV cache across two, four, or eight devices, buying memory headroom at the price of an all-reduce on every layer and a hardware bill that grows with the device count. The algorithms community shrinks the cache in place, with KV quantisation and eviction keeping a single GPU and spending a little quality instead. Compression papers report memory ratios, parallel-scaling papers report throughput curves, and almost nobody puts the two on the same cost axis. We place tensor-parallel configurations (degree 1 to 8) and KV-compressed configurations (16/8/4-bit, keep-ratios down to 0.25) on one costnormalised axis, cost per million tokens against latency, using a profiled simulator calibrated on A100, A40, and H100 hardware, and we go looking for the cost-equivalence crossover. We do not find one. Across two models (Llama-2 at 7B and 70B), three GPU types, and every level of memory relief we could construct, compression is cheaper by 1.20x to 2.00x. A 7B model on an 80 GB device cannot exhaust its KV budget within its own context window, and the boundary that decides between the strategies is model size relative to device memory, at roughly 36B parameters for an 80 GB card. Below that wall, compression dominates and extra GPUs are largely wasted spend; above it, tensor parallelism stops being a choice and becomes an entry ticket: Llama-2-70B is infeasible on one A100 at any KV setting, because the binding resource is weights, which KV compression does not touch. Tensor parallelism is the only lever that improves latency (compression makes per-token latency worse, by 8 to 93%, through batching contention), while compression is the only lever that multiplies capacity per dollar (16.5x, against 1.21x for an eightfold spend on GPUs).