arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.19296cs.AR

HyperCut:基于有向超图与早期过滤的快速层间调度

HyperCut: Fast Inter-Layer Scheduling via Directed Hypergraph and Early Filtering

Ziang Wei, Zirui Xu, Sufeng Guo, Chuanchao Gao, Yiyang Gao, Arvind Easwaran, Yuxiang Fu

首次发表
浏览论文内容

中文总结 AI 辅助

HyperCut 是基于有向超图的分层划分映射框架,通过早期过滤优化层间调度,相比 SET 性能提升 2.0 倍、探索时间减少 80.47%,设计空间从 O(9.899^N) 缩小至 O(N)。

中文摘要 AI 辅助

随着深度神经网络(DNN)规模不断扩大,层间调度(即协调计算资源的空间分配与跨层的时间执行顺序)已成为在 tiled 加速器上维持高利用率和能效的决定性因素。然而,现有的层间调度器会等到完成完整的细粒度层内调度后才反馈代价,这种解耦流程会反复探索次优甚至不可行的层间调度方案,且层间阶段缺乏早期剪枝是 DNN 编译器设计空间探索(DSE)的关键瓶颈。我们的核心发现是:一旦层间切割确定了子网格形状,层内调度的代价就可以被紧密地上界,这使得我们无需求解层内问题就能评估每个层间候选方案。因此,我们提出了一种分层划分与映射框架 HyperCut,该框架可基于超图划分对层间调度进行早期过滤。基于 DNN 的有向超图(DHG)抽象,我们引入了统一表示 State,其共同编码 DHG 划分、 tile 网格分配和张量批量拆分,从而将划分与映射耦合为一个联合优化对象。对于含 N 层的 DNN,由此产生的理论设计空间被限定为 O(N),而最先进的开源调度器 SET 的设计空间为 O(9.899^N)。在 10 个评估案例中,以几何均值衡量,HyperCut 相比 SET 基准实现了 2.0 倍的性能提升和 80.47% 的探索时间减少。

英文摘要

As deep neural networks (DNNs) continue to scale, inter-layer scheduling, which orchestrates the spatial allocation of compute resources and the temporal execution order across layers, has become a decisive factor in sustaining high utilization and energy efficiency on tiled accelerators. However, existing inter-layer schedulers defer cost feedback until a complete fine-grained intra-layer scheduling has been resolved. The resulting decoupled flow repeatedly explores sub-optimal or even infeasible inter-layer schedules, and the absence of early pruning during the inter-layer phase remains a critical bottleneck for design-space exploration (DSE) in DNN compilers. Our key observation is that the cost of an intra-layer scheduling can be tightly upper-bounded once the inter-layer cut fixes the sub-mesh shape, which lets us cost every inter-layer candidate without solving the intra-layer problem. Hence, we propose a hierarchical partitioning-and-mapping framework, HyperCut, that enables early filtering of inter-layer schedules based on hypergraph partitioning. Based on the directed hypergraph (DHG) abstraction of DNN, we introduce a unified representation, State, that jointly encodes the DHG partition, tile mesh allocation and tensor batch splitting. Thereby, partitioning and mapping are coupled into a union optimization object. For a DNN with N layers, the resulting theoretical design space is bounded by O(N), compared with O(9.899^N) for the state-of-the-art open-source scheduler SET. Across 10 evaluated cases, HyperCut achieves 2.0x performance improvement and 80.47% exploration time reduction over the SET baseline, measured by geometric mean.

补充信息

↑