发表机构
Dell Technologies(戴尔科技集团)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出跨部署栈的AI推理优化统一分析方法,引入三层分类法,将部署建模为多目标优化,提出证据协议提升基准可比性,并综合边缘与数据中心平台证据,揭示跨层交互决定部署结果。
AI 中文摘要
AI部署性能不仅由模型架构单独决定,还受压缩、编译器转换和服务策略之间交互作用的影响。已发表的基准测试通常在不可比较的条件下报告延迟和吞吐量,限制了其在部署决策中的使用。本文提出了一种跨部署栈的推理优化统一分析方法。我们引入了一个三层分类法,涵盖模型级技术(如量化、剪枝和蒸馏)、编译器转换(如图融合、布局优化和内核自动调优)以及系统策略(如动态批处理、准入控制和内存分层)。我们将部署问题表述为在准确性、延迟、吞吐量、内存占用和能耗上的约束多目标优化问题,并分析了一个具有帕累托单调性和尺度不变性的部署排序泛函。Roofline模型展示了内存带宽层次如何在不同精度范围内限制性能,而排队模型解释了服务时间变化如何在负载下放大响应时间。为提高可比性,我们提出了一种证据协议,该协议区分了测量、派生和分析性声明;将数值比较限制在论文内部结果;并要求报告硬件、软件版本、批处理语义和热状态。我们综合了来自边缘平台(包括Jetson AGX Orin和五个推理框架)、数据中心GPU(包括A100和H100及三个LLM服务引擎)以及Llama-3.1系列量化研究的证据。综合结果表明,部署结果由跨层交互决定,任何单层分析都无法预测。最后,我们提出了一个约束感知的选择程序,并讨论了编译器-服务协同优化、跨硬件性能预测和标准化能耗报告中的开放问题。
英文摘要
AI deployment performance is shaped not by model architecture alone, but by interactions among compression, compiler transformations, and serving policies. Published benchmarks often report latency and throughput under incomparable conditions, limiting their use for deployment decisions. This paper presents a unified analytical treatment of inference optimization across the deployment stack. We introduce a three-layer taxonomy covering model-level techniques such as quantization, pruning, and distillation; compiler transformations such as graph fusion, layout optimization, and kernel autotuning; and system policies such as dynamic batching, admission control, and memory tiering. We formulate deployment as a constrained multi-objective optimization problem over accuracy, latency, throughput, memory footprint, and energy, and analyze a deployment-ranking functional with Pareto monotonicity and scale invariance. Roofline models show how memory-bandwidth hierarchies bound performance across precision regimes, while queuing models explain how service-time changes amplify response time under load. To improve comparability, we propose an evidence protocol that separates measured, derived, and analytical claims; limits numerical comparison to within-paper results; and requires reporting of hardware, software versions, batch semantics, and thermal state. We synthesize evidence from edge platforms, including Jetson AGX Orin and five inference frameworks; data center GPUs, including A100 and H100 with three LLM serving engines; and quantization studies across the Llama-3.1 family. The synthesis shows that deployment outcomes are governed by cross-layer interactions that no single-layer analysis can predict. We conclude with a constraint-aware selection procedure and open problems in compiler-serving co-optimization, cross-hardware performance prediction, and standardized energy reporting.
Comments29 pages, 12 figures