arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

面向移动异构推理协同执行的分区感知调度

Partition-Aware Scheduling for Mobile Heterogeneous Inference Co-Execution

Zhuojin Li, Marco Paolieri, Leana Golubchik

arXiv 2609.14213首次发表:更新:

发表机构

University of Southern California(南加州大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出一种分区感知的DAG调度框架,联合优化算子分区、设备分配与执行顺序,在移动异构平台上实现接近离线最优的推理延迟,且调度开销极低。

AI 中文摘要

现代移动推理运行在结合移动GPU与多个CPU核心簇的异构平台上。现有优化通常利用算子间并行性(通过将整个算子分配给CPU核心或GPU)或算子内并行性(通过将每个算子分区以进行CPU-GPU协同执行)。我们综合考虑这两种形式的并行性,以改善可由具有预定义输入/输出张量形状的静态算子DAG(例如CNN或视觉Transformer)表示的任务的推理延迟。我们定义了移动异构推理的分区感知DAG调度问题,说明最佳策略取决于推理DAG的结构,从而促使我们提出一个联合公式,以捕获算子分区选择、设备分配和执行顺序。我们提出了一个在线迭代搜索框架,该框架将大型DAG分解为阶段,将搜索集中在关键算子上,并使用延迟预测器来估计分区执行,而无需进行详尽的性能剖析。在代表性的移动推理工作负载中,我们的方法实现了接近离线解决方案的延迟,同时将调度开销保持在模型初始化成本的一小部分,从而允许在部署时进行平台特定的调度。

英文摘要

Modern mobile inference runs on heterogeneous platforms combining mobile GPUs with multiple CPU core clusters. Existing optimizations typically exploit either inter-operator parallelism, by assigning entire operators to CPU cores or to the GPU, or intra-operator parallelism, by partitioning each operator for CPU-GPU co-execution. We consider these two forms of parallelism together, to improve inference latency of tasks that can be represented by a static DAG of operators with predefined input/output tensor shapes (e.g., CNNs or vision transformers). We define the problem of partition-aware DAG scheduling for mobile heterogeneous inference, illustrating that the best strategy depends on the structure of the inference DAG, thus motivating a joint formulation capturing operator partition choices, device assignment, and execution order. We propose an online iterative search framework, which decomposes large DAGs into stages, focuses search on critical operators, and uses latency predictors to estimate partitioned execution without exhaustive profiling. Across representative mobile inference workloads, our approach achieves latency close to an offline solution while keeping scheduling overhead to a fraction of the model initialization cost, allowing platform-specific scheduling at deployment time.

CommentsAccepted for publication in the Performance Evaluation journal (presented at IFIP Performance 2026, in Ghent, Belgium)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑