发表机构
MIT; NVIDIA(麻省理工学院; 英伟达)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
AccelForge是一个统一的AI加速器建模与协同设计框架,通过可组合模型、快速映射器和高性能Python实现,显著提升评估速度与易用性。
AI 中文摘要
张量代数工作负载(其中深度神经网络是典型代表)是现代数据中心和边缘部署中能耗密集的工作负载,这使得加速器成为实现能效和高吞吐量的必要条件。为了快速评估和迭代加速器设计,我们需要一个加速器建模框架,该框架能够捕捉设备、电路、架构和工作负载的关键属性,并优化工作负载到硬件的映射。在本文中,我们介绍了AccelForge,它在能力、速度和易用性方面改进了现有的加速器建模框架。AccelForge将多项工作统一到一个框架中,并包括:(1)可组合的用户定义和用户可修改的设备、电路和架构模型,(2)快速映射器,能够在数量级更少(计算机和人工)时间内实现准确评估,以及(3)易于使用且易于扩展但仍具有高性能的模型和映射器的Python实现,以支持快速研究和对新颖优化的扩展。
英文摘要
Tensor algebra workloads, of which deep neural networks are prominent examples, are energy-intensive workloads in modern datacenter and edge deployments, making accelerators necessary to achieve energy efficiency and high throughput. To quickly evaluate and iterate on accelerator designs, we need an accelerator modeling framework that captures salient attributes of devices, circuits, architectures, workloads, as well as optimizing the mapping of the workload onto the hardware. In this paper, we introduce AccelForge, which improves upon existing accelerator modeling frameworks in capabilities, speed, and ease-of-use. AccelForge unifies and multiple works into one framework, and it includes (1) composable user-defined and user-modifiable models of devices, circuits, and architectures, (2) fast mappers that enable accurate evaluation in orders of magnitude less (computer and human) time, and (3) easy-to-use and easy-to-extend, yet still high performance, Python implementations of both the model and mapper to enable rapid research and extension to novel optimizations.