AI 中文总结
研究针对异构平台上LLM推理,提出硬件感知的DOPS框架,通过构建阶段感知DAG,集成双焦点调度器和权重布局仲裁器,实现算子调度与权重布局联合优化,在异构系统中取得加速效果并支持相关分析。
AI 中文摘要
预填充-解码分解(PD)和基于屋顶线的算子放置是在异构系统中划分大语言模型(LLM)推理的常见策略,但实际中往往不足。端到端延迟还取决于工作负载形状、运行时设备争用和持久权重布局。我们提出了DOPS(动态算子调度),这是一个硬件感知的闭环框架,可联合优化算子调度和逐块权重布局。DOPS构建了一个阶段感知有向无环图(DAG),并集成了两个组件:用于动态算子到设备放置的双焦点调度器和用于在严格内存约束下选择硬件高效权重布局的权重布局仲裁器(WLA)。在结合神经处理单元(NPUs)和内存处理(PIM)设备的代表性异构系统中,双焦点比PD基线实现了1.20倍至2.23倍的几何平均加速。WLA比双焦点/线性提供了1.28倍至1.33倍的额外几何平均加速。DOPS还支持对LLM服务的工作负载敏感性和硬件可扩展性进行系统分析。
英文摘要
Prefill-decode disaggregation (PD) and roofline-based operator placement are common strategies for partitioning Large Language Model (LLM) inference across heterogeneous systems, but they are often insufficient in practice. End-to-end latency also depends on workload shape, runtime device contention, and persistent weight layout. We present DOPS (dynamic operator scheduling), a hardware-aware, closed-loop framework that jointly optimizes operator scheduling and blockwise weight layouts. DOPS constructs a stage-aware directed acyclic graph (DAG) and integrates two components: the Bifocal scheduler for dynamic operator-to-device placement and the Weight Layout Arbiter (WLA) for selecting hardware-efficient weight layouts under strict memory constraints. Across representative heterogeneous systems combining neural processing units (NPUs) and processing-in-memory (PIM) devices, Bifocal achieves geometric-mean speedups of 1.20$\times$ to 2.23$\times$ over the PD baseline. WLA provides an additional geometric-mean speedup of 1.28$\times$ to 1.33$\times$ over Bifocal/Linear. DOPS also supports systematic analysis of workload sensitivity and hardware scalability for LLM serving. The source code is available at https://github.com/YIAI-02/TriForm, and the visualization tool is demonstrated at https://youtu.be/Ya_oMCyYno0.
CommentsTo appear in MICRO 2026