人工智能解码芯片的经济学:重新平衡计算、容量和带宽以实现高效的大语言模型推理
The Economics of AI Decoding Chips: Rebalancing Compute, Capacity, and Bandwidth for Efficient LLM Inference
浏览论文内容
中文总结 AI 辅助
研究大语言模型解码中GPU计算与内存不匹配问题,提出用F/B和F/S常数衡量效率,主张重新平衡解码加速器,介绍Skymizer HTX-301优势,如成本低、无受限输入,能以较低成本实现高效推理。
中文摘要 AI 辅助
每款主流GPU都构建为计算能力强而容量弱:其算术吞吐量巨大,但内存却太少,无法容纳现代模型。相比之下,大语言模型解码所需计算少但内存量大:GPU的浮点单元在解码期间利用率仅为个位数百分比,且所需内存只能与更多计算捆绑销售。只有在超大规模下,即专家混合(MoE)模型分布在96至320个GPU的专家并行集群中为数千个并发用户服务时,计算能力才能得到恢复,而这种规模只有少数运营商能够达到。我们用两个固定的每芯片常数形式化了这种低效率。F/B(屋顶线脊点)决定计算能力能否被利用;F/S(每GB内存捆绑的计算量)决定必须购买多少计算能力。然后我们主张一种重新平衡的解码加速器:计算能力更少、更多的通用内存以及故意更低且更便宜的带宽。专用的28nm PCIe加速器Skymizer HTX-301占据了该设计点。其入门成本低。一张由八块芯片组成的卡可容纳约19,000美元的DeepSeek-R1 671B,一台由四张四芯片卡组成的4U服务器可为两个用户提供服务,每个用户每秒确定性地处理20.3个令牌,成本约为28,000美元。两者成本均低于单个H100,而该模型的最低GPU部署是一个近350,000美元的八GPU节点。通过添加硬件,并发能力可以扩展:八台4U服务器可为十六个用户提供服务,成本约为224,000美元,为节点价格的三分之二,每百万令牌成本约为12美元,而节点为21美元。HTX-301的决定性优势是供应链中没有任何受限输入:它不使用高带宽内存、不使用CoWoS且不使用前沿逻辑。
英文摘要
Every mainstream GPU is built compute-heavy and capacity-light: it pairs enormous arithmetic throughput with too little memory to hold a modern model. In contrast, large language model decoding requires little compute and a large amount of memory: a GPU's floating-point units run at single-digit-percent utilization during decoding, and the memory the workload does need is sold only bundled with yet more compute. The compute is recovered only at hyperscale, where Mixture-of-Experts (MoE) models are spread across 96--320-GPU expert-parallel clusters serving thousands of concurrent users, a scale available to a handful of operators. We formalize the inefficiency with two fixed per-chip constants. F/B, the roofline ridge point, determines whether the compute can be utilized; F/S, the compute bundled with each GB of memory, determines how much compute must be bought. We then argue for a rebalanced decode accelerator: less compute, far more commodity memory, and a deliberately lower and cheaper bandwidth. The Skymizer HTX-301, a purpose-built 28nm PCIe accelerator using commodity DDR5, occupies that design point. Its entry cost is low. A single eight-chip card holds DeepSeek-R1 671B for about \$19,000, and a 4U server of four four-chip cards serves two users at a deterministic 20.3 tokens per second each for about \$28,000. Either costs less than a single H100, while the minimum GPU deployment for the model is an eight-GPU node near \$350,000. Concurrency then scales out by adding hardware: eight 4U servers carry sixteen users for about \$224,000, two-thirds of the node's price, with the cost per token unchanged at about \$12 per million against the node's \$21. The HTX-301's decisive advantage is a supply chain free of every rationed input: it uses no high-bandwidth memory, no CoWoS, and no leading-edge logic.