arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SynapticOS:资源受限微控制器上神经处理单元的推理优先运行时架构

SynapticOS: An Inference-First Runtime Architecture for Neural Processing Units on Resource-Constrained Microcontrollers

Dimitrios Kafetzis

arXiv 2607.12606首次发表:更新:

AI 中文总结

针对资源受限微控制器上神经处理单元系统软件问题,提出基于Zephyr的开源运行时SynapticOS,含四个协作子系统,经评估给出构建占用、推理时间等数据,61测试套件在CI模拟器路径通过率达100%,且以Apache 2.0发布。

AI 中文摘要

片上神经处理单元(NPU)的微控制器已成为主流,但承载它们的系统软件却并非如此。像Zephyr或FreeRTOS与TensorFlow Lite Micro的生产组合将AI推理视为应用层库,导致内存碎片化、加速器状态维护及模型生命周期保护成为应用开发者反复面临的问题。我们展示了SynapticOS的第一阶段基础,这是一个基于Zephyr构建的开源运行时,将推理视为一等工作负载。它有四个协作子系统:张量感知的内存块分配器,具有16字节DMA对齐的持久和临时生命周期,共享单个区域,零碎片化且恒定时间分配;NPU和DSP的四态硬件抽象层;三态模型生命周期注册表;四标记周期精确剖析器。我们在NXP FRDM-MCXN947和qemu_cortex_m3模拟器上进行评估,给出了不同的构建占用空间和端到端推理时间等数据,还有一个61测试套件在CI模拟器路径上的测试结果。SynapticOS以Apache 2.0发布。

英文摘要

Microcontrollers with on-die neural processing units (NPUs) have become mainstream, but the system software hosting them has not: production combinations of Zephyr or FreeRTOS with TensorFlow Lite Micro treat AI inference as an application-layer library, leaving memory fragmentation, accelerator-state hygiene, and model-lifecycle guards as recurring application-developer concerns. We present the Phase 1 foundation of SynapticOS, an open-source runtime built on Zephyr that treats inference as a first-class workload. It contributes four cooperating subsystems: (1) a tensor-aware bump allocator with 16-byte DMA-aligned persistent and ephemeral lifetimes sharing a single arena, achieving constant-time allocation (~154 cycles per call, ~78,000 allocations per second at 150 MHz, invariant across tensor sizes) with zero fragmentation by construction; (2) a four-state hardware abstraction layer for the NPU and DSP, implemented by a deterministic software stub (for CI under QEMU) and a Neutron-flavoured backend (for the NXP MCXN947); (3) a three-state model lifecycle registry with duplicate-name detection, idempotent load/unload, and hot-swap guards; and (4) a four-mark cycle-accurate profiler. We evaluate on the NXP FRDM-MCXN947 (dual Cortex-M33 at 150 MHz) and the qemu_cortex_m3 emulator. Build footprints are 67 KB flash / 184 KB SRAM on FRDM (shell, 128 KB arena) and 24 KB flash / 28 KB SRAM on QEMU (no shell, 8 KB arena). End-to-end inference brackets through the deterministic stub kernel measure 1,038 us on FRDM and 781 us on QEMU for a 16x16x3 INT8 input; these are baseline overhead numbers, not Neutron silicon measurements, which arrive with the real SDK invoke path in Phase 2. A 61-test suite across 10 ZTEST suites passes 100% in 6.6 s on the CI emulator path. SynapticOS is released under Apache 2.0 at https://github.com/Dimitrios-Kafetzis/SynapticOS

Comments12 pages, 5 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑