arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

MeshReduce-U:面向Mesh NoC上不规则神经归约的编译器引导式通信缩减方案

MeshReduce-U: Compiler-Guided Communication Reduction for Irregular Neural Reductions on Mesh NoCs

Amirreza Khorasanian

arXiv 2608.26220首次发表:更新:

AI 中文总结

MeshReduce-U是面向Mesh NoC空间加速器的编译器引导式通信缩减与路由框架,通过结构重写减少神经归约通信开销,在各类工作负载上显著降低延迟与链路利用率。

AI 中文摘要

许多不规则神经工作负载会引发带有重复邻域和非局部通信的倾斜多对一归约。传统NoC映射器会针对固定通信图优化布局与路由,不过关联归约提供了在路由前消除流量的合法机会。本文提出MeshReduce-U,一种基于Mesh NoC的空间加速器所用的编译器引导式通信缩减与路由框架。MeshReduce-U会合并共置源节点、形成本地聚合岛、阻塞具有兼容扇入结构的通道、选择容量可行的汇节点,并使用融合的感知资源利用率代价路由剩余固定宽度载体。确定性路由重放模型会分别报告由调度推导得到的通信延迟、总链路利用率(TLU)和融合链路利用率(FusedTLU)。在包含20种可降低神经网络工作负载的数据集上,MeshReduce-U相较于ABC风格的源聚合基线,分别将平均延迟、TLU和FusedTLU降低了40.3%、56.0%和48.7%,在所有工作负载上均改善了这三个指标;在40种合成不规则归约上,其分别将平均延迟和TLU降低了12.3%和19.7%。一项包含30个实例的逐次传递研究进一步表明,该结构重写操作将全局载体数量减少了60.9%,重放延迟降低了63.0%。这些结果说明,在路由前对可归约神经通信进行重写,比在未缩减的流量图上进行更深入的搜索更为有效。

英文摘要

Many irregular neural workloads induce skewed many-to-one reductions with repeated neighborhoods and nonlocal communication. Conventional NoC mappers optimize placement and routes for a fixed communication graph, even though associative reductions expose legal opportunities to eliminate traffic before routing. We present MeshReduce-U, a compiler-guided communication-reduction and routing framework for mesh-NoC-based spatial accelerators. MeshReduce-U coalesces colocated sources, forms local aggregation islands, blocks channels with compatible fan-in structure, selects capacity-feasible sinks, and routes the remaining fixed-width carriers using fused usage-aware costs. A deterministic route-replay model reports schedule-derived communication latency, total link usage (TLU), and fused link usage (FusedTLU) separately. Across a 20-workload lowerable neural-network zoo, MeshReduce-U reduces mean latency, TLU, and FusedTLU by 40.3%, 56.0%, and 48.7%, respectively, relative to an ABC-style source-aggregation baseline, improving all three metrics on every workload. Across 40 synthetic irregular reductions, it reduces mean latency and TLU by 12.3% and 19.7%. A new 30-instance pass-by-pass study further shows that the structural rewrites reduce the global carrier count by 60.9% and replay latency by 63.0%. These results show that rewriting reducible neural communication before routing can be more effective than searching harder over an unreduced traffic graph.

Comments19 pages, 7 figures. Extended preprint

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑