arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.05420cs.DC

在单个 NVIDIA L20 上持续运行 70B 级 AWQ 推理:吞吐量、稳定性、能耗与质量表征

Sustained 70B-Class AWQ Inference on a Single NVIDIA L20: Throughput, Stability, Energy, and Quality Characterization

Yin Li

首次发表
浏览论文内容

中文总结 AI 辅助

本报告在单个 NVIDIA L20 48GB GPU 上使用 vLLM 和 AWQ Marlin 服务 Qwen2.5-72B-Instruct-AWQ,通过 24 小时浸泡测试验证了其可持续的吞吐量、稳定性、能耗与质量,证明该配置可作为面向吞吐量的 70B 级端点。

中文摘要 AI 辅助

服务 70B 级开放权重语言模型通常与 80GB 加速器、张量并行多 GPU 系统或供应商管理的推理配置文件相关联。本技术报告评估单个 NVIDIA L20 48GB GPU 能否持续承载有用的 70B 级量化服务工作负载。我们测量了在单个 L20 上使用 vLLM 0.8.5.post1 和 AWQ Marlin 服务的 Qwen2.5-72B-Instruct-AWQ。在约 512 个输入 token 和 256 个输出 token 的固定工作负载下,24 小时并发数为 10 的浸泡测试完成了 36,740/36,740 个请求,无请求失败,无 vLLM CUDA 内存不足、回溯或进程被杀死的迹象。系统维持了 108.84 输出 token/秒的吞吐量,p95 首 token 时间为 6.61 秒,p95 端到端延迟为 23.54 秒。通过 nvidia-smi 采样的 GPU 板级功耗在运行期间估计为 7.92 千瓦时,对应 0.330 输出 token/焦耳和 1.008 总 token/焦耳。在并发数为 1、4、8 和 16 的情况下重复固定形状运行,12/12 次运行全部成功;并发数为 16 的条件在三次运行中平均输出吞吐量为 127.22 ± 12.68 输出 token/秒。同一 AWQ 端点在 MMLU 上取得 0.8130 的绝对质量分数,在 CMMLU 上为 0.8309,在 GSM8K 上为 0.8082,此外还完成了 80/80 的 MT-Bench 答案生成和 60 项的 8K LongBench 子集。证据支持一个狭窄的论断:在测试的固定形状工作负载下,经过精心配置的单个 L20 可以将 Qwen2.5-72B-Instruct-AWQ 作为面向吞吐量的 70B 级端点来服务。这并不证明无损的 AWQ 质量保持、低延迟交互式服务、广泛的生产 SLA 覆盖,或与 BF16/FP16 基线的等价性。

英文摘要

Serving 70B-class open-weight language models is usually associated with 80GB accelerators, tensor-parallel multi-GPU systems, or vendor-managed inference profiles. This technical report evaluates whether a single NVIDIA L20 48GB GPU can sustain a useful 70B-class quantized serving workload. We measure Qwen2.5-72B-Instruct-AWQ served with vLLM 0.8.5.post1 and AWQ Marlin on one L20. Under a fixed workload of approximately 512 input tokens and 256 output tokens, a 24-hour concurrency-10 soak completed 36,740/36,740 requests with no request failures and no vLLM CUDA OOM, traceback, or killed-process signatures. The system sustained 108.84 output tokens/s, with p95 time-to-first-token of 6.61s and p95 end-to-end latency of 23.54s. GPU-board power sampled through nvidia-smi produced an estimated 7.92 kWh over the run, corresponding to 0.330 output tokens/J and 1.008 total tokens/J. Repeated fixed-shape runs at concurrency 1, 4, 8, and 16 completed 12/12 runs successfully; the concurrency-16 condition averaged 127.22 +/- 12.68 output tokens/s over three runs. The same AWQ endpoint also produced absolute quality scores of 0.8130 on MMLU, 0.8309 on CMMLU, and 0.8082 on GSM8K, plus 80/80 MT-Bench answer generations and a 60-item 8K LongBench subset. The evidence supports a narrow claim: a carefully configured single L20 can serve Qwen2.5-72B-Instruct-AWQ as a throughput-oriented 70B-class endpoint under the tested fixed-shape workload. It does not prove lossless AWQ quality retention, low-latency interactive serving, broad production SLA coverage, or equivalence to a BF16/FP16 baseline.

↑