发表机构
Mozilla(Mozilla)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究在英特尔 TDX 下 NVIDIA H100 GPU 上,比较标准与机密计算模式对语言模型推理的影响。通过两个模型实验,测量多项指标,发现机密模式会使首次令牌时间增加、全局令牌吞吐量下降,较大模型更早饱和,为容量规划提供参考。
AI 中文摘要
机密计算正成为处理敏感输入或保护专有模型资产的人工智能推理工作负载的实际部署要求。然而,为 GPU 加速的大语言模型服务启用机密执行的性能成本仍取决于工作负载且在操作上很重要。本文进行了一项基准研究,在英特尔 TDX 机密实例中的单个 NVIDIA H100 80GB GPU 上比较标准非机密执行和机密计算模式。评估使用两个代表性语言模型,测量了首次令牌时间、端到端请求延迟等指标。在固定请求率实验中,机密模式使 Mistral-7B 的平均首次令牌时间增加 21.8%,Qwen3-30B-A3B 增加 27.8%,同时全局令牌吞吐量分别下降 17.7%和 21.1%。在闭环并发实验中,吞吐量差距在 11.5 - 20.2%范围内,但较大模型在机密模式下更早达到饱和拐点。结果表明,机密 GPU 推理在负载下可保持可用吞吐量,但容量规划必须考虑稳定的吞吐量损失和较大模型更早的饱和行为。
英文摘要
Confidential computing is becoming a practical deployment requirement for AI inference workloads that process sensitive inputs or protect proprietary model assets. However, the performance cost of enabling confidential execution for GPU-accelerated large language model serving remains workload dependent and operationally important. This paper presents a benchmark study comparing standard non-confidential execution with confidential computing mode on a single NVIDIA H100 80GB GPU hosted in an Intel TDX confidential instance. The evaluation uses two representative language models, Mistral-7B v0.1 and Qwen3-30B-A3B, and measures time to first token, end-to-end request latency, per-request token generation throughput, global token throughput, and closed-loop request throughput under increasing concurrency. In fixed request-rate experiments, confidential mode increases average TTFT by 21.8% for Mistral-7B and 27.8% for Qwen3-30B-A3B, while global token throughput drops by 17.7% and 21.1%, respectively. In closed-loop concurrency experiments, throughput gaps remain in the 11.5-20.2% range, but the larger model reaches its saturation knee earlier under confidential mode. The results suggest that confidential GPU inference can retain usable throughput under load, but capacity planning must account for both the steady throughput penalty and the earlier saturation behavior observed for larger models.