arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.11939cs.ARcs.CVcs.DC

自适应AI:边缘智能视觉系统上的能效多出口TinyML

Adaptive AI: Energy Efficient Multi-exit TinyML on Intelligent Vision Systems at the Edge

Luca Crupi, Lorenzo Lamberti, Alessandro Giusti, Daniele Palossi

首次发表
浏览论文内容

中文总结 AI 辅助

本研究提出一种基于置信度门控的多出口TinyML方案,在GAP9 SoC上实现MobileNetV2动态推理,相比标准模型降低41%计算成本、29%推理时间和24%能耗,精度仅损失约1%。

中文摘要 AI 辅助

传统的边缘设备TinyML系统通过依赖固定深度模型来实现高精度,这些模型无论输入复杂度如何,都需要恒定数量的乘加(MAC)运算。这种方法在电池供电的物联网(IoT)设备中浪费了关键资源,并限制了边缘信息物理系统的实时性能。多出口执行方案缓解了这些问题,并广泛用于高端设备(如GPU),但在边缘物联网设备上很少被利用,因为考虑到其严格的内存和计算约束,需要大量的重新思考。我们通过设计并在超低功耗GWT GAP9片上系统(SoC)上部署一种新颖的多出口计算方案来解决这些方面,并在用于ImageNet-100分类任务的MobileNetV2卷积神经网络(CNN)上进行了演示。我们的方法在CNN的不同深度引入了多个出口,每个出口都有基于置信度的门控机制,动态且自主地决定是继续还是停止推理。将我们的多出口策略与GAP9 SoC上的标准MobileNetV2进行比较,我们展示了平均计算成本降低41%(从313 MMAC降至185 MMAC),推理时间降低29%(从49毫秒降至35毫秒),以及每帧能量节省24%(从2.1毫焦降至1.6毫焦)。所有这些改进都伴随着与全深度MobileNetV2相比约1%的精度损失,后者达到了80.5%的精度。最后,将我们的自适应多出口方案与同样部署在GAP9上的第三方最先进自适应CNN进行比较,我们实现了其计算效率的两倍以上,从8.1 MAC/周期提高到17.2 MAC/周期。

英文摘要

Traditional TinyML systems for edge devices achieve high accuracy by relying on fixed-depth models that require a constant number of multiply-accumulate (MAC) operations regardless of the input complexity. This approach wastes critical resources in battery-powered Internet-of-Things (IoT) devices and limits the real-time performance of edge cyber-physical systems. Multi-exit execution schemes mitigate these issues and are widely used on high-end devices such as GPUs, but are rarely exploited on edge IoT devices because they require substantial rethinking given their strict memory and computational constraints. We address these aspects by designing and deploying, on an ultra-low-power GWT GAP9 System-on-Chip (SoC), a novel multi-exit computational scheme, demonstrating it on a MobileNetV2 convolutional neural network (CNN) for the ImageNet-100 classification task. Our approach introduces multiple exits at different CNN depths, each with a confidence-based gating mechanism that dynamically and autonomously decides whether to continue or stop inference. Comparing our multi-exit strategy to the standard MobileNetV2 on a GAP9 SoC, we show a 41% reduction in the average computational cost (from 313 MMAC to 185 MMAC), a 29% lower inference time (from 49 to 35 ms), and an energy saving of 24% (from 2.1 to 1.6 mJ per frame). All these improvements come with a ~1% loss in accuracy compared to the full-depth MobileNetV2, which achieves 80.5%. Finally, comparing our adaptable multi-exit scheme with a third-party state-of-the-art adaptive CNN, also deployed on the GAP9, we achieve more than 2x its computational efficiency, increasing it from 8.1 to 17.2 MAC/cycle.

发表机构

  • IDSIA(IDSIA研究所)
  • SUPSI(瑞士南部应用科学与艺术大学)

机构由 AI 辅助整理,请以论文原文为准。

↑