arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.10805cs.CVcs.AI

基于I/O感知重构的快速且内存高效的小波卷积

Fast and Memory-Efficient Wavelet Convolutions via I/O-Aware Reformulation

发表机构内盖夫本-古里安大学
查看机构详情
  • Ben-Gurion University of the Negev(内盖夫本-古里安大学)

机构由 AI 辅助整理,请以论文原文为准。

Amit Aflalo, Shahaf E. Finder, Roy Amoyal, Eran Treister, Oren Freifeld

首次发表
浏览论文内容

中文总结 AI 辅助

该研究针对WTConv受内存限制的问题,通过三种代数重构实现I/O感知融合实现,大幅减少HBM流量,提升训练速度并降低内存占用,保留WTConv优势的同时消除其系统开销。

中文摘要 AI 辅助

小波卷积(WTConv)已成为标准卷积的热门替代方案,它能随分解层数指数级扩展网络感受野,同时保持参数量线性增长。然而其参考实现因通过高带宽内存(HBM)的过度数据移动而严重受内存限制。我们构建WTConv的I/O模型以表征此瓶颈,并据此指导三种代数重构:(1)在芯片上重新计算低成本的Haar分析蝶形运算;(2)将多级合成级联折叠为按输出坐标位索引的单个闭式传递;(3)将学习到的逐通道尺度折叠进卷积权重。这些重构共同实现了I/O感知的融合实现,大幅减少HBM流量。我们在不同分解层数和广泛的张量形状上评估WTConvNeXt配置,尽管算术操作相当,但参考WTConv比其替代的深度卷积慢得多。我们的重构将建模的HBM流量减少约2.55倍,相比参考实现最高实现4.35倍的训练加速,同时峰值内存使用量大致减半。因此,我们的重构保留了WTConv的优势,同时大幅降低其执行时间和内存占用,消除了此前限制其实际效率的系统开销。

英文摘要

Wavelet convolution (WTConv) has emerged as an increasingly popular drop-in replacement for standard convolutions, expanding a network's receptive field exponentially with the number of decomposition levels while keeping the parameter count linear. However, its reference implementation is severely memory-bound due to excessive data movement through high-bandwidth memory (HBM). We develop an I/O model of WTConv to characterize this bottleneck and use it to guide three algebraic reformulations: (1) recomputing the inexpensive Haar analysis butterfly on chip, (2) collapsing the multi-level synthesis cascade into a single closed-form pass indexed by output-coordinate bits, and (3) folding learned per-channel scales into the convolution weights. Together, these reformulations enable an I/O-aware fused implementation that substantially reduces HBM traffic. We evaluate the WTConvNeXt configuration across decomposition levels and a broad range of tensor shapes. Despite performing comparable arithmetic, the reference WTConv is substantially slower than the depthwise convolution it replaces. Our reformulation reduces modeled HBM traffic by approximately $2.55\times$, yielding up to a $4.35\times$ training speedup over the reference while roughly halving peak memory usage. Thus, our reformulation preserves the benefits of WTConv while substantially reducing its execution time and memory footprint, removing the systems overhead that previously limited its practical efficiency.

↑