AI 中文总结
本文在NVIDIA B200 GPU上结合Intel TDX机密VM与NVIDIA Blackwell GPU的CC技术,测试大语言模型推理/训练在TEE中的性能,得出正确配置下机密推理吞吐量仅增1%-3%等结果并给出部署指南。
AI 中文摘要
本文在NVIDIA B200 GPU上的可信执行环境(TEE)中运行大语言模型推理和训练,测量其性能影响,采用的是英特尔信任域扩展(TDX)机密虚拟机,搭配NVIDIA Blackwell GPU上的机密计算(CC)技术。性能影响通过单台物理主机上的机密运行与非机密运行的配对对比得出,唯一变量是GPU CC位和虚拟机启动中的TDX客体对象。主要结果显示,当堆栈配置正确时,Blackwell上的机密推理吞吐量开销仅为低个位数,约1%-3%;而默认推理堆栈因可避免的配置而非可达到的运行点,会产生30%-40%的性能损失。该成本无法用单一数值完全体现,因为它受两个独立维度支配:一是随批次大小增长而摊销的每主机操作固定成本,二是跟踪加密集体通信步骤耗时占比的每NVLink流量成本,两者中哪一个占主导取决于工作负载和软件。我们将每项成本定位到特定的加密边界,提供了一个微基准测试,可将服务惩罚预测到提交计数范围内,最后给出具体的部署指南。GPU计算、能耗和可用内存容量不受CC影响。
英文摘要
This paper measures the performance impact of running large language model inference and training inside a Trusted Execution Environment (TEE) on NVIDIA B200 GPUs, using Intel Trust Domain Extensions (TDX) confidential VMs together with NVIDIA Confidential Computing (CC) on Blackwell GPUs. The performance impact is derived from paired confidential versus non-confidential runs on a single physical host where the only variable is the GPU CC bit and the TDX guest object in the VM launch. The main result is that confidential inference on Blackwell achieves low single-digit throughput overhead when the stack is configured correctly, at about 1-3%. Stock inference stacks incur 30 to 40% penalties due to avoidable configurations rather than the achievable operating point. The cost is not fully represented by a single number because it is governed by two independent axes, a fixed per-host-operation cost that amortizes as batch size grows and a per-NVLink-traffic cost that tracks the share of the step spent in encrypted collectives, and which of the two dominates is set by the workload and the software. We localize each cost to a specific encrypted boundary, give a microbenchmark that predicts the serving penalty to within a submission count, and end with concrete deployment guidance. GPU compute, energy draw, and usable memory capacity are unaffected by CC.
Comments23 pages, 3 figures, 22 tables