DAN调度器:通用NPUs调度、内存布局和流水线重叠的确定性三阶段协同优化
LATTICE: Constraint-Directed Scheduling, Memory Planning, and Pipeline Refinement for NPUs
浏览论文内容
中文总结 AI 辅助
研究针对通用NPUs,提出DAN-Scheduler框架,通过内存压力感知拓扑调度、确定性线性重新打包、关键路径增强三个阶段协同优化调度、内存布局和流水线重叠,在六个运算符级DAG上评估,相比基线有显著性能提升。
中文摘要 AI 辅助
神经处理单元(NPUs)越来越多地用于高吞吐量、内存受限的推理,但其分层片上内存以及异构计算和数据移动引擎紧密耦合了执行顺序、内存布局和流水线重叠。现有编译器流程通常分别优化这些维度,导致片上驻留过多、不必要的片外流量以及资源未充分利用。本文提出了DAN-Scheduler,一种用于内核内NPU执行的确定性离线调度和编译器优化框架。它分三个阶段协同优化这些决策。内存压力感知拓扑调度(MPAS)重新排序运算符以缩短张量生命周期并减少峰值片上内存使用。确定性线性重新打包(DLR)构建无冲突内存布局并应用分层感知、成本感知溢出启发式算法以减少有限容量下的碎片化和片外流量。关键路径增强(CPE)在保持前两个阶段建立的内存行为的同时改善计算-DMA重叠。我们在从实际达芬奇NPU收集并在通用NPU执行模型上重放的六个跟踪派生的运算符级DAG上评估了DAN-Scheduler。与四个强大的外部基线相比,DAN-Scheduler在所有24个工作负载指标单元上取得了最佳或并列最佳结果,平均而言,相对于最佳外部竞争对手,峰值内存、额外DDR流量、溢出计数和完工时间分别减少了18.3%、20.4%、14.2%和16.3%。相对于原始调度,它在相同指标上分别减少了38.3%、62.0%、64.9%和57.5%。这些结果表明,确定性的阶段式协同优化对于内存受限的NPU执行是有效的。代码和数据可在指定网址获取。
英文摘要
General-purpose NPUs execute fine-grained command DAGs across heterogeneous compute and memory-transfer engines backed by finite, explicitly managed on-chip memories. This execution model creates a directed dependency between scheduling and memory planning: different legal topological orders induce different lifetime overlap, placement opportunities, and spill behavior, while a materialized layout introduces physical-address reuse constraints absent from the input precedence DAG. Command order therefore shapes the feasible memory plan, and the realized plan in turn defines the legal space for subsequent timing refinement. We present LATTICE, a deterministic constraint-directed compiler pipeline. Memory-Pressure-Aware Topological Scheduling reshapes lifetime geometry before address binding; Deterministic Linear Repackaging materializes tiered placement, spill/reload events, and plan-induced reuse constraints; and Critical Path Enhancement recovers pipeline parallelism while preserving the selected memory plan. Every accepted schedule passes independent memory and timing verification. Across six artifact-provided command traces labeled as derived from a Da Vinci NPU flow, LATTICE achieves the best or tied-best result in all 24 evaluated workload-metric comparisons. Relative to the best evaluated baseline for each workload and metric, it reduces peak memory, extra DDR traffic, spill count, and modeled makespan by 18.3% lower, 20.4% lower, 14.1% lower, and 16.3% lower, respectively. Plan-preserving CPE further reduces makespan by 12.1% over Freeze while leaving placement and memory traffic unchanged, establishing the static memory plan as a verifiable scheduling contract between memory planning and pipeline optimization.