发表机构
National Cheng Kung University; Qualcomm Technologies, Inc.(国立成功大学; 高通技术有限公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出一种 LLM 智能体驱动的方法,将 PyTorch 模型自动转换为多个异构推理运行时(如 OpenVINO、RKNN、TensorRT、ONNX Runtime)的可执行推理,通过分阶段验证和领域知识注入,实现 FP16 精度部署,并提供了多运行时部署的工程实践经验。
AI 中文摘要
边缘 AI 模型部署是一个多阶段的工程过程,涉及模型转换、算子兼容性处理、运行时集成和精度验证。虽然先前的工作已在 Qualcomm AI Runtime 上展示了基于智能体的自动化,但更广泛的边缘推理运行时生态系统,包括 Intel OpenVINO、Rockchip RKNN、NVIDIA TensorRT 和 ONNX Runtime,呈现出不同的工具链和优化策略。本文扩展了 AIPC(AI 移植转换,一种先前在 Qualcomm AI Runtime 上演示的 LLM 智能体驱动的 AI 模型部署自动化方法)到多运行时场景,提出了一种 LLM 智能体驱动的方法,用于跨异构推理后端(如 Intel OpenVINO、Rockchip RKNN、NVIDIA TensorRT 和 ONNX Runtime)的自动化单模型到单运行时部署。我们将边缘 AI 部署分解为标准化、可验证的阶段,并通过智能体技能、辅助脚本和分阶段验证循环将特定于运行时的领域知识注入智能体执行流程。使用代表性的视觉模型,我们证明了基于智能体的部署可以完成从 PyTorch 模型到其可执行推理的转换,针对 x86/NPU 的 OpenVINO、RK3588 的 RKNN、NVIDIA GPU 的 TensorRT 以及 Qualcomm NPU 的 ONNX Runtime,并专注于 FP16 精度部署可行性验证。本文的贡献主要在于提供多运行时部署工程实践经验、工具链映射分析、一个布局适配和推理替换层,该层将手动转置插入从智能体的修复负担中移除,以及在结构化知识注入下智能体偏差行为的经验表征,而非大规模系统基准测试或跨运行时算子修复策略比较。
英文摘要
Edge AI model deployment is a multi-stage engineering process involving model conversion, operator compatibility handling, runtime integration, and precision verification. While prior work has demonstrated agent-based automation for Qualcomm AI Runtime, the broader edge inference runtime ecosystem, including Intel OpenVINO, Rockchip RKNN, NVIDIA TensorRT, and ONNX Runtime, presents distinct toolchains and optimization strategies. This paper extends AIPC (AI Porting Conversion, an LLM agent-driven methodology for AI model deployment automation previously demonstrated on Qualcomm AI Runtime) to multi-runtime scenarios, proposing an LLM agent-driven approach for automated single-model-to-single-runtime deployment across heterogeneous inference backends, such as Intel OpenVINO, Rockchip RKNN, NVIDIA TensorRT, and ONNX Runtime. We decompose the edge AI deployment into standardized, verifiable stages, and inject runtime-specific domain knowledge into the agent execution flow through agent skills, auxiliary scripts, and staged verification loops. Using representative vision models, we demonstrate that agent-based deployment can complete the conversion from a PyTorch model to its executable inference, targeting OpenVINO for x86/NPU, RKNN for RK3588, TensorRT for NVIDIA GPU, and ONNX Runtime for Qualcomm NPU with a focus on FP16 precision deployment feasibility verification. The contributions of this paper primarily lie in providing multi-runtime deployment engineering practice experience, toolchain mapping analysis, a layout-adaptation and inference-replacement layer that removes manual transpose insertion from the agent's repair burden, and an empirical characterization of agent deviation behavior under structured knowledge injection, rather than large-scale systematic benchmarking or cross-runtime operator repair strategy comparison.
Comments13 pages with 8 figures and 7 tables. Prepared with ACM conference format