arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.12614cs.OS

作为操作系统对象的推理管道:微控制器神经推理的优先级调度和恒定占用流

Inference Pipelines as Operating-System Objects: Priority Scheduling and Constant-Footprint Streaming for Microcontroller Neural Inference

Dimitrios Kafetzis

首次发表
浏览论文内容

中文总结 AI 辅助

研究针对微控制器神经推理中推理管道的处理问题,提出将其作为操作系统对象的方法,通过优先级调度等技术,实现恒定占用流,评估了在不同平台上的性能,弥补了阶段差距,虽未达加速目标,但测试通过,有一定贡献。

中文摘要 AI 辅助

微控制器运行时将推理管道(预处理、加速器调用、后处理)视为应用程序代码:每个项目都围绕库调用重新实现阶段排序、缓冲区大小调整和完成信号。我们认为这些是操作系统的关注点,并展示了SynapticOS的第二阶段推理引擎,这是一个基于开源Zephyr的运行时,使管道成为一流的操作系统对象。管道从静态池中提取,根据规范阶段顺序进行验证,并由优先级作业调度器(实时>正常>尽力而为,每个类按FIFO)执行,具有取消功能和有界作业表;推理路径上没有堆。阶段缓冲区根据九个内置处理器的配置和张量几何形状精确调整大小(用户阶段有4倍的有界回退);所有中间结果都存放在每帧重置的临时区域中,因此流占用是恒定的。我们在NXP FRDM-MCXN947(Cortex-M33,150 MHz)和qemu_cortex_m3 CI目标上进行评估,两者都运行确定性存根NPU内核:是引擎开销基线而非芯片吞吐量。在板上,调度器比第一阶段直接HAL括号增加了92微秒(1130对1038微秒;调度1微秒);一个30帧、六阶段的面部检测管道平均每帧4.63毫秒(215.8 FPS,包括存根模型),而在QEMU软浮点下为31.1毫秒,在恒定的2784字节区域峰值下每帧返回零。PowerQuad DSP针对FFT和Q15矩阵乘法进行了路由和自校准;端到端加速比为5.51倍(256点FFT)和1.66倍(16x16矩阵乘法),未达到计划的10倍目标——报告为未达到,而非重新调整范围。阶段边界分析现在可以在板上实时运行,弥补了第一阶段的差距。该引擎在QEMU上增加了3.8 KB闪存,在FRDM上增加了20.7 KB。在13个ZTEST套件中的99个测试在模拟下100%通过。以Apache 2.0协议在这个https URL发布。

英文摘要

Microcontroller runtimes treat the inference pipeline -- pre-processing, accelerator invocation, post-processing -- as application code: every project re-implements stage sequencing, buffer sizing, and completion signalling around a library call. We argue these are operating-system concerns and present the Phase 2 inference engine of SynapticOS, an open-source Zephyr-based runtime that makes the pipeline a first-class OS object. A pipeline is drawn from a static pool, validated against a canonical stage order, and executed by a priority job scheduler (realtime > normal > best-effort, FIFO per class) with cancellation and a bounded job table; no heap on the inference path. Stage buffers are sized exactly from configuration and tensor geometry for the nine built-in processors (bounded 4x fallback for user stages); all intermediates live in an ephemeral arena reset per frame, so streaming footprint is constant. We evaluate on the NXP FRDM-MCXN947 (Cortex-M33, 150 MHz) and the qemu_cortex_m3 CI target, both running a deterministic stub NPU kernel: engine-overhead baselines, not silicon throughput. On the board the scheduler adds 92 us over the Phase 1 direct-HAL bracket (1,130 vs 1,038 us; dispatch 1 us); a 30-frame, six-stage face-detection pipeline averages 4.63 ms/frame (215.8 FPS, stub model included) vs 31.1 ms under QEMU soft-float, at a constant 2,784-byte arena peak returning to zero each frame. The PowerQuad DSP is routed and self-calibrated for FFT and Q15 matmul; end-to-end speedups are 5.51x (256-point FFT) and 1.66x (16x16 matmul), short of the plan's 10x target -- reported as missed, not re-scoped. Stage-boundary profiling now runs live on the board, closing a Phase 1 gap. The engine adds 3.8 KB flash on QEMU and 20.7 KB on FRDM. 99 tests across 13 ZTEST suites pass 100% under emulation. Released under Apache 2.0 at https://github.com/Dimitrios-Kafetzis/SynapticOS

补充信息

↑