MEDEM:多引擎深度学习加速器设计方法
MEDEM: Multi-Engine DL Accelerator Design Methodology
浏览论文内容
中文总结 AI 辅助
MEDEM提出多引擎深度学习加速器系统化设计方法,通过引擎抽象与协同设计高效搜索指数级设计空间,在多种资源预算下实现EDP最高4.84倍、吞吐量1.59倍的提升。
中文摘要 AI 辅助
多引擎深度学习(DL)加速器正变得越来越普遍,因为它们解决了现代深度学习工作负载的异构性和日益增长的复杂性。为了高效处理多样化的深度学习工作负载,这些加速器必须结合具有互补能力的引擎组合,以匹配这些工作负载中异构内核的不同计算特性。然而,现有的多引擎深度学习加速器设计方法缺乏系统性的方法论,留下了基本问题未解决。这些问题包括如何为具有不同计算特性的工作负载协同设计引擎,以及哪些引擎组合能最小化这些工作负载的总执行成本(如时间或能量)。解决这些问题需要对指数级大的设计空间进行高效探索。为了系统地解决这些问题,本工作提出了多引擎深度学习加速器设计方法(MEDEM)。MEDEM定义了通用的引擎抽象,协同设计候选实例(引擎),并选择协同设计引擎的组合,以在给定的多样化深度学习工作负载和资源预算下最小化总执行成本。MEDEM包含一套设计策略,能高效地在引擎协同设计和组合选择的指数级大设计空间中导航,识别出高度优化的多引擎加速器。综合评估表明,MEDEM识别出的加速器优于最先进的设计,在能量延迟积(EDP)方面实现了高达4.84倍的几何平均改进,在吞吐量方面实现了1.59倍的改进。这些改进是在不同的资源预算下实现的,展示了MEDEM的可扩展性,并使用51个单模型和多模型深度学习工作负载,展示了其泛化能力。
英文摘要
Multi-engine deep learning (DL) accelerators are becoming increasingly prevalent as they address the heterogeneity and growing complexity of modern DL workloads. To efficiently process diverse DL workloads, these accelerators must incorporate combinations of engines with complementary capabilities to match the distinct computational characteristics of these workloads' heterogeneous kernels. However, existing multi-engine DL accelerator design approaches lack a systematic methodology, leaving fundamental questions unresolved. These include how to co-design engines for workloads with diverse computational characteristics and which engine combinations minimize aggregate execution costs (such as time or energy) across such workloads. Addressing these questions requires efficient exploration of exponentially large design spaces. To address these questions systematically, this work proposes Multi-Engine DL Accelerator Design Methodology (MEDEM). MEDEM defines generic engine abstractions, co-designs candidate instances (engines), and selects a combination of co-designed engines to minimize aggregate execution cost given diverse DL workloads and a resource budget. MEDEM encompasses a set of design strategies that efficiently navigate the exponentially large design spaces of engine co-design and combination selection, identifying highly optimized multi-engine accelerators. A comprehensive evaluation demonstrates that MEDEM identifies accelerators that outperform state-of-the-art designs, delivering geometric-mean improvements of up to 4.84x in energy-delay product (EDP) and 1.59x in throughput. The improvements are achieved using different resource budgets, demonstrating MEDEM's scalability, and using 51 single- and multi-model DL workloads, demonstrating its generalizability.
发表机构
- Chalmers University of Technology and University of Gothenburg(查尔姆斯理工大学和哥德堡大学)
机构由 AI 辅助整理,请以论文原文为准。