arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

HeteroMosaic:为高效能边缘大语言模型推理揭示并利用异构执行机会

HeteroMosaic: Exposing and Exploiting Heterogeneous Execution Opportunities for Energy-Efficient Edge LLM Inference

Gregory Hyegang Jun, Wesley Pang, Eddie Richter, Mehdi Saeedi, Aporva Amarnath, Pallavi Ferrao, Deming Chen

arXiv 2607.12839首次发表:更新:

AI 中文总结

研究针对边缘大语言模型推理中异构资源利用不足问题,提出HeteroMosaic框架。通过异构屋顶线模型确定组合iGPU和NPU执行的时机,分解推理为微批次并进行跟踪引导协同优化。在多平台评估中取得显著加速和能耗降低,性能优于现有方案。

AI 中文摘要

现代边缘系统级芯片(SoC)集成了CPU、集成GPU(iGPU)和神经处理单元(NPU),但现有大语言模型运行时通常仅做粗略的设备级决策或单独优化算子。这导致异构资源未被充分利用。本文提出HeteroMosaic,一种针对边缘大语言模型推理的优先考虑异构性的调度框架。它先用异构屋顶线模型确定何时组合iGPU和NPU执行有益,再将推理分解为保留依赖的微批次以暴露跨加速器重叠,并在内存争用、DVFS、设备差异和NPU运行时开销等实际影响下进行调度和设备分配的跟踪引导协同优化。在三个AMD Ryzen AI平台上评估,在平衡平台上,HeteroMosaic比iGPU基线加速高达1.73倍,比NPU基线加速1.78倍,比其他框架加速2.05倍,同时能耗降低达45.3%,性能比先前异构边缘AI解决方案提升达2.35倍。

英文摘要

Modern edge system-on-chips (SoCs) combine CPUs, integrated GPUs (iGPUs), and neural processing units (NPUs), yet existing LLM runtimes typically make coarse device-level decisions or optimize operators in isolation. As a result, they underutilize heterogeneous resources, particularly on unified-memory platforms where performance depends on both device placement and task-graph coordination. We present HeteroMosaic, a heterogeneity-first scheduling framework for edge LLM inference. HeteroMosaic first uses a heterogeneous roofline model to identify when combining iGPU and NPU execution is beneficial. It then decomposes inference into dependency-preserving micro-batches that expose cross-accelerator overlap and applies trace-guided co-optimization of scheduling and device allocation under practical effects such as memory contention, DVFS, device variation, and NPU runtime overheads. We implement HeteroMosaic in PyTorch C++ and evaluate it on three AMD Ryzen AI platforms spanning NPU-heavy, balanced, and iGPU-heavy designs. On the balanced platform, HeteroMosaic achieves up to 1.73X speedup over an iGPU baseline, 1.78X over an NPU baseline, and 2.05X over frameworks such as llama dot cpp, while reducing energy by up to 45.3%. It also improves performance over prior heterogeneous edge AI solutions by up to 2.35X.

CommentsAccepted at the 59th IEEE/ACM International Symposium on Microarchitecture (MICRO 2026)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑