arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.11954cs.LG

高效AI模型部署:基于量化分析工具

Efficient AI Model Deployment Using Quantization Analysis Tool

Dwith Chenna, Kanishka Macherla

首次发表
浏览论文内容

中文总结 AI 辅助

本文介绍一种基于ONNX的量化分析工具,通过逐层敏感性分析与分布可视化,指导精度选择,在保持准确率的同时优化模型大小和延迟,实现高效AI模型部署。

中文摘要 AI 辅助

随着深度学习模型越来越多地部署在资源受限的设备上,对高效模型优化技术的需求持续增长。在边缘和低功耗平台上有效部署AI模型,需要采用能够减小模型规模、降低计算成本同时保持高准确率的优化方法。本文介绍了量化分析工具(Quantization Analysis Tool),这是一个旨在简化量化工作流程并支持性能高效模型部署的实用系统。该工具基于ONNX框架构建,以实现广泛的互操作性,提供详细的逐层敏感性分析、权重和激活分布的可视化,以及指导精度选择的洞察。通过识别对降低精度具有弹性或敏感性的层,该工具使开发者能够在模型大小、延迟和准确率之间做出明智的权衡。跨多种神经网络架构的实验评估表明,该工具能有效提升量化后的准确率,从而在实际部署场景中提高效率。该工具还为开发者提供了关于量化对模型及其准确率影响的宝贵洞察。这项工作凸显了该工具的能力、实际应用,以及其通过稳健的量化分析实现高效AI模型部署的作用。

英文摘要

As deep learning models are increasingly deployed on resource constrained devices, the demand for efficient model optimization techniques continues to grow. Effective deployment of AI models on edge and low power platforms requires optimization methods that reduce model size and computational cost while maintaining high accuracy. This paper presents Quantization Analysis Tool, a practical system designed to streamline quantization workflows and support performance efficient model deployment. Built on the ONNX framework for broad interoperability, the tool provides detailed layer-wise sensitivity analysis, visualization of weight and activation distributions, and insights to guide precision selection. By identifying layers that are resilient or sensitive to reduced precision, the tool enables developers to make informed trade-offs between model size, latency, and accuracy. Experimental evaluations across multiple neural network architectures demonstrate that the tool effectively improves the quantized accuracy, leading to improved efficiency in real-world deployment scenarios. The tool also provides developers valuable insights into the effects on quantization on the model and its accuracy. This work highlights the tools capabilities, practical applications, and its role in enabling efficient AI model deployment through robust quantization analysis

发表机构

  • University of Maryland(马里兰大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑