发表机构
The Ohio State University(俄亥俄州立大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文调研近期边缘AI部署研究,提炼指南并在多平台测试,发现不同任务适用不同压缩技术,剪枝可能提升分割性能但会增加延迟,相关成果已开源。
AI 中文摘要
在资源受限的边缘设备上运行大型AI模型需要模型压缩以减小模型体积和计算量。然而,压缩效果好的技术未必部署效果好。我们调研了数十篇报告了真实硬件压缩结果的近期研究,并从中提炼出实用的部署指南。遵循这些指南,我们在GPU、CPU和Raspberry Pi平台上部署了适用于问答和图像分割任务的轻量化语言模型与图像模型。没有单一技术能在所有任务中胜出:对于问答任务,Qwen3.5 0.8B在Q5_K_M GGUF量化下达到93.85的SQuAD F1值和92的EM值,而在1%比例下的结构化剪枝则会损失16个F1值;对于分割任务,排序结果相反:默认量化不会改变参数和MACs,而剪枝则能在mIoU几乎不变的情况下将模型体积削减近80%。剪枝甚至可能因破坏k-quant超级块对齐而使部署产物体积膨胀21-49%,结合更长、格式合规性更低的输出,这会使Raspberry Pi的延迟最高提升3.4倍。压缩还可能产生能力的表象而非明显的性能下降:一个LoRA恢复的变体保持完全可解析,在将100个预测中的97个归类到单一类别的情况下,仍保持71%的严格BoolQ准确率和52.6%的平衡准确率。我们通过神经流图分析和预填充-解码级延迟分解解释这些效应,并将其凝练为任务特定的部署研究方向。合适的技术取决于任务、模型和硬件。我们的实验代码和产物已开源至该https URL。
英文摘要
Running large AI models on resource-constrained edge devices requires model compression to reduce model size and computation. What compresses well, however, need not deploy well. We survey dozens of recent works that report compression results on real hardware and extract practical deployment guidelines from them. Following these guidelines, we deploy compact language and image models on GPU, CPU, and Raspberry Pi platforms across question answering and image segmentation. No single technique wins across tasks. For question answering, Qwen3.5 0.8B reaches 93.85 SQuAD F1 and 92 EM under Q5_K_M GGUF quantization, while structured pruning at the same precision costs 16 F1 at a 1% ratio. For segmentation, the ranking reverses: default quantization leaves parameters and MACs unchanged, whereas pruning cuts model size by nearly 80% at near-constant mIoU. Pruning can even inflate the deployed artifact by 21-49% by breaking k-quant super-block alignment; combined with longer, less format-compliant outputs, this raises Raspberry Pi latency up to 3.4x. Compression can also manufacture the appearance of competence rather than destroy it visibly: one LoRA-recovered variant stays fully parseable and holds 71% strict BoolQ accuracy while sending 97 of 100 predictions to a single class, at 52.6% balanced accuracy. We explain these effects through neural-flow graph analysis and prefill-decode-level latency decomposition, and condense them into task-specific deployment research directions. The right technique depends on the task, the model, and the hardware. Our experiment code and artifacts are open-sourced at https://github.com/Arnavvvkumar/deployment
CommentsParts of this work were presented at the IEEE Consumer Communications & Networking Conference (CCNC), Las Vegas, NV, USA, January 2026