发表机构
University College Dublin(都柏林大学学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对ViT在边缘设备部署的存储与自适应问题,提出运行时自适应流水线,将预训练ViT转为单二进制,支持动态计算切换,实现4.86倍存储缩减和2.8倍加速。
AI 中文摘要
在低功耗边缘设备上部署视觉Transformer(ViT)因高计算需求而颇具挑战性。传统剪枝框架需要为每个稀疏度级别单独编译二进制文件,增加了存储开销并限制了运行时适应性。本文提出了一种端到端部署流水线,将预训练的ViT转换为单个运行时可配置的二进制文件,从而在嵌入式CPU上实现动态计算预算切换。这是通过重构生成的C内核,修改循环边界和二进制掩码控制逻辑来实现的,允许执行通过紧凑的外部配置文件在不同离散稀疏度级别之间切换。与多二进制部署相比,所提出的运行时自适应方法将设备上存储减少了高达4.86倍,ViT-Base仅需163 MB,而原先接近800 MB。为了最大化剪枝效率,我们为多层感知器(MLP)层引入了一种硬件对齐的块剪枝策略。此外,提出了一种自定义ISA扩展,以利用线性投影内核中的输入重用模式。在Synopsys TRV32P3FX RISC-V处理器上,对于ViT-Base,在65% MLP和50%注意力头剪枝下,整个系统实现了高达2.8倍的加速。仅ISA扩展就提供了1.56倍的加速和33%的推理能耗降低,在TSMC 28 nm实现中面积开销为24.7%。
英文摘要
Deploying Vision Transformers (ViTs) on low-power edge devices is challenging due to high computational demands. Conventional pruning frameworks require a separate compiled binary for each sparsity level, increasing storage overhead and limiting runtime adaptability. This paper presents an end-to-end deployment pipeline that transforms pretrained ViTs into a single runtime-configurable binary, enabling dynamic compute-budget switching on embedded CPUs. This is achieved by restructuring generated C kernels with modified loop bounds and binary-mask control logic, allowing execution to switch across discrete sparsity levels via compact external configuration files. Compared to multi-binary deployment, the proposed runtime-adaptive approach reduces on-device storage by up to 4.86x, requiring only 163 MB for ViT-Base instead of nearly 800 MB. To maximize pruning efficiency, we introduce a hardware-aligned block pruning strategy for Multi-Layer Perceptron (MLP) layers. In addition, a custom ISA extension is proposed to exploit input-reuse patterns in linear projection kernels. On a Synopsys TRV32P3FX RISC-V processor, the full system achieves up to 2.8x speedup at 65% MLP and 50% attention-head pruning for ViT-Base. The ISA extension alone provides a 1.56x speedup and 33% lower inference energy, with a 24.7% area overhead in a TSMC 28 nm implementation.
CommentsAccepted for publication at the 2026 IEEE 33rd International Conference on Electronics, Circuits and Systems (ICECS)