发表机构
DEVCOM Army Research Laboratory; University of West Florida(DEVCOM陆军研究实验室; 西佛罗里达大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出与硬件和模型无关的广义优化引擎(GOE)架构,用于加速边缘AI推理,经GOE压缩的语言模型可在无GPU的边缘CPU部署运行,压缩方法选择影响部署后的任务精度。
AI 中文摘要
人工智能(AI)模型在多个领域展现出卓越能力,但其广泛部署受限于显著的计算成本,尤其在资源受限设备上。本文探究各类AI模型优化技术、算法与抽象的理论基础,讨论其降低计算复杂度、内存占用、延迟及功耗的潜力。此外,我们提出一种与硬件(HW)和模型无关的综合广义优化架构,整合这些技术以提升效率。本研究强调此类广义优化系统在战术环境中为资源受限异构硬件准备模型部署的关键作用。作为具体演示,我们表明经GOE压缩的语言模型可在无GPU的边缘CPU上部署并运行,且压缩方法的选择(而非仅其标称位宽)决定任务精度能否在部署后保留。
英文摘要
Artificial intelligence (AI) models have demonstrated remarkable capabilities across various domains, yet their widespread deployment is impeded by significant computational costs, particularly on resource-constrained devices. This paper explores the theoretical underpinnings of various AI model optimization techniques, algorithms, and abstractions, discussing their potential to reduce computational complexity, memory footprint, latency, and power consumption. Furthermore, we propose a comprehensive hardware (HW) and model-agnostic generalized optimization architecture that integrates these techniques for improved efficiency. Our study underscores the critical role of such a generalized optimization system in preparing model deployment over resource-constrained heterogeneous hardware in a tactical environment. As a concrete demonstration, we show that GOE-compressed language models deploy and run on a GPU-less edge CPU, and that the choice of compression method, not merely its nominal bit-width, determines whether task accuracy survives deployment.
Comments14 pages, compressed version is under review in IEEE MILCOM 2026