arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

原位右对齐(RiP)卷积:一种简单、通用且近乎最优的CNN推理内存高效策略

Right In-Place (RiP) Convolution: A Simple, General, and Near-Optimal Strategy for Memory-Efficient CNN Inference

Opegbemi Matthias Busoye, Tolulope Matthew Busoye, Eghonghon-aye Eigbe

arXiv 2610.00586首次发表:更新:

发表机构

PowerLabs Technologies(PowerLabs Technologies)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对CNN推理中的激活内存瓶颈,提出原位右对齐(RiP)卷积,修正了现有最优公式的缺陷,推广至任意卷积参数,在保持位完全相同输出的同时平均节省24.8%内存,并提升微控制器上可部署模型数量。

AI 中文摘要

激活内存,而非计算,限制了卷积神经网络(CNN)在微控制器等受限硬件上的推理。直接原位卷积消除了双缓冲区的开销,但Gural和Murmann提出的内存最优公式假设了有效填充、单位步长、单位膨胀和奇数方形卷积核,并且需要非顺序遍历,导致转置操作使推理时间增加2倍。我们发现了他们发表的闭式解不成立的两种情况:(1)恰好有$(k-1)C_{in} \bmod (C_{out}-C_{in})$个标量的分配不足,这在他们自己部署的网络的每个卷积层上都会出现,表现为对仍存活的输入的静默损坏;(2)一旦关键路径离开输出网格,会出现高达$2{,}432$倍的无界高估。我们修正了这两个问题,并将其推广到任意步长、膨胀、填充和矩形卷积核。然后,我们提出了原位右对齐(RiP)卷积,这是一种位完全相同的操作,其中每一层在共享工作空间中右对齐读取其输入,并从索引零开始左对齐写入其输出。债务是输出像素索引的分段仿射函数,因此以$O(1)$复杂度评估其断点即可获得最小安全间隙,而无需枚举输出网格,同时保持行主序访问。在$10{,}000$个随机层上,RiP未产生任何损坏;在来自25种架构的84个卷积层上,它与人字形工作空间在58个层上完全匹配,在81个层上误差在5%以内,平均比双缓冲少使用24.8%的内存。将其写入TinyEngine的内核并部署到Raspberry Pi Pico 1和Pico 2上,在11个MCUNet模型上,峰值激活内存减少了12.5%至33.3%,同时保持周期数和位完全相同输出不变,使得适合Pico 1的256 KB SRAM的模型数量从六个增加到九个。

英文摘要

Activation memory, not compute, limits CNN inference on constrained hardware such as microcontrollers. Direct in-place convolution removes the dual-buffer cost, but the memory-optimal formulation of Gural and Murmann assumes valid padding, unit stride, unit dilation, and odd square kernels, and needs a non-sequential traversal costing $2\times$ inference time in transposes. We identify two regimes in which their published closed form does not hold: (1) an under-allocation of exactly $(k-1)C_{in} \bmod (C_{out}-C_{in})$ scalars, active on every convolutional layer of their own deployed network and manifesting as a silent corruption of still-live input; (2) an unbounded overestimate, up to $2{,}432\times$, once the critical leg leaves the output grid. We correct both and generalize to arbitrary stride, dilation, padding, and rectangular kernels. We then propose Right In-Place (RiP) convolution, a bit-identical operation in which every layer reads its input right-aligned in a shared workspace and writes its output left-aligned from index zero. The debt is piecewise affine in the output pixel index, so evaluating its breakpoints in $O(1)$ yields the minimum safe gap without enumerating the output grid, with row-major access preserved. Across $10{,}000$ random layers RiP produced no corruption, and across 84 convolutional layers from 25 architectures it matches the herringbone workspace exactly on 58 and within 5% on 81, using 24.8% less memory than dual buffering on average. Written into TinyEngine's kernels and deployed to a Raspberry Pi Pico 1 and Pico 2, it cuts peak activation memory across eleven MCUNet models by 12.5 to 33.3% at unchanged cycle counts and bit-identical outputs, raising the number of models that fit the Pico 1's 256 KB SRAM from six to nine.

CommentsExtended version of a paper accepted at the NeurIPS 2026 Workshop on Global South in AI

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑