HyperCut:基于有向超图与早期过滤的快速层间调度
HyperCut: Fast Inter-Layer Scheduling via Directed Hypergraph and Early Filtering
浏览论文内容
中文总结 AI 辅助
HyperCut 是基于有向超图的分层划分映射框架,通过早期过滤优化层间调度,相比 SET 性能提升 2.0 倍、探索时间减少 80.47%,设计空间从 O(9.899^N) 缩小至 O(N)。
中文摘要 AI 辅助
随着深度神经网络(DNN)规模不断扩大,层间调度(即协调计算资源的空间分配与跨层的时间执行顺序)已成为在 tiled 加速器上维持高利用率和能效的决定性因素。然而,现有的层间调度器会等到完成完整的细粒度层内调度后才反馈代价,这种解耦流程会反复探索次优甚至不可行的层间调度方案,且层间阶段缺乏早期剪枝是 DNN 编译器设计空间探索(DSE)的关键瓶颈。我们的核心发现是:一旦层间切割确定了子网格形状,层内调度的代价就可以被紧密地上界,这使得我们无需求解层内问题就能评估每个层间候选方案。因此,我们提出了一种分层划分与映射框架 HyperCut,该框架可基于超图划分对层间调度进行早期过滤。基于 DNN 的有向超图(DHG)抽象,我们引入了统一表示 State,其共同编码 DHG 划分、 tile 网格分配和张量批量拆分,从而将划分与映射耦合为一个联合优化对象。对于含 N 层的 DNN,由此产生的理论设计空间被限定为 O(N),而最先进的开源调度器 SET 的设计空间为 O(9.899^N)。在 10 个评估案例中,以几何均值衡量,HyperCut 相比 SET 基准实现了 2.0 倍的性能提升和 80.47% 的探索时间减少。
英文摘要
As deep neural networks (DNNs) continue to scale, inter-layer scheduling, which orchestrates the spatial allocation of compute resources and the temporal execution order across layers, has become a decisive factor in sustaining high utilization and energy efficiency on tiled accelerators. However, existing inter-layer schedulers defer cost feedback until a complete fine-grained intra-layer scheduling has been resolved. The resulting decoupled flow repeatedly explores sub-optimal or even infeasible inter-layer schedules, and the absence of early pruning during the inter-layer phase remains a critical bottleneck for design-space exploration (DSE) in DNN compilers. Our key observation is that the cost of an intra-layer scheduling can be tightly upper-bounded once the inter-layer cut fixes the sub-mesh shape, which lets us cost every inter-layer candidate without solving the intra-layer problem. Hence, we propose a hierarchical partitioning-and-mapping framework, HyperCut, that enables early filtering of inter-layer schedules based on hypergraph partitioning. Based on the directed hypergraph (DHG) abstraction of DNN, we introduce a unified representation, State, that jointly encodes the DHG partition, tile mesh allocation and tensor batch splitting. Thereby, partitioning and mapping are coupled into a union optimization object. For a DNN with N layers, the resulting theoretical design space is bounded by O(N), compared with O(9.899^N) for the state-of-the-art open-source scheduler SET. Across 10 evaluated cases, HyperCut achieves 2.0x performance improvement and 80.47% exploration time reduction over the SET baseline, measured by geometric mean.