Para-Pipe:在片上系统(SoC)上利用机器学习计算图的分层算子并行性
Para-Pipe: Exploiting Hierarchical Operator Parallelism of ML Computational Graphs on SoCs
浏览论文内容
中文总结 AI 辅助
Para-Pipe是整合算子并行性的分层映射框架,可平衡SoC上深度学习的吞吐量与延迟,提升能效,在两款SoC上分别实现11.0%和23.3%的能效提升。
中文摘要 AI 辅助
随着基于边缘的深度学习应用日益复杂,在异构片上系统(SoC)上优化性能面临独特挑战。传统流水线技术将计算分布在不同片上处理单元间,虽对吞吐量有效,但未解决现代神经网络因复杂依赖关系和广泛算子并行性带来的延迟需求。利用算子并行性实现多处理单元并发执行以降低推理延迟具有潜力,但优先选择流水线或并行执行常需权衡,优化某一性能指标会对另一指标产生不利影响。本文提出Para-Pipe,一种在流水线架构内集成阶段内和阶段间算子并行性的分层映射框架。Para-Pipe通过选择性微调流水线阶段内部及阶段间的并行度水平,在吞吐量与延迟间权衡,此策略可显著降低处理器间通信开销,大幅提升能效。我们的评估表明,Para-Pipe生成多个帕累托最优配置,在配备ARM CPU和GPU的晶晨半导体(Amlogic)SoC,以及带有深度学习加速器和两个数字信号处理器(DSP)的黑芝麻科技(Black Sesame Technology)SoC上实现了吞吐量与延迟的平衡。更重要的是,在晶晨半导体SoC上,Para-Pipe下吞吐量优化配置的平均能效较纯流水线策略提升11.0%,较非流水线并行执行提升23.3%。
英文摘要
As edge-based deep learning applications become more complex, optimizing performance on heterogeneous System-on-Chips (SoCs) presents unique challenges. Traditional pipelining techniques distributing the computation across different on-chip processing units, while effective for throughput, do not address the latency demands posed by modern neural networks with complex interdependencies and extensive operator parallelism. There is a potential in leveraging operator parallelism to enable concurrent execution across multiple processing units, thereby reducing inference latency. However, prioritizing pipelining or parallel execution often necessitates a compromise, where optimizing one performance metric adversely impacts the other. This paper introduces Para-Pipe, a hierarchical mapping framework that integrates intra- and inter-stage operator parallelism within a pipelined architecture. Para-Pipe navigates the trade-off between throughput and latency by selectively fine-tuning parallelism levels within and across pipeline stages. This strategy can significantly reduce inter-processor communication overhead, significantly improving energy efficiency. Our evaluation demonstrates that Para-Pipe generates multiple Pareto-optimal configurations, achieving a balance between throughput and latency on an Amlogic SoC equipped with ARM big.LITTLE CPUs and GPU, as well as the Black Sesame Technology SoC featuring a deep learning accelerator and two DSPs. More importantly, throughput-optimized configurations under Para-Pipe on Amlogic SoC show an average energy efficiency improvement of 11.0% over purely pipelined strategies and 23.3% relative to non-pipelined parallel execution.
发表机构
- School of Computing, National University of Singapore(新加坡国立大学计算机学院)
- the University of Amsterdam(阿姆斯特丹大学)
- Black Sesame Technologies(黑芝麻智能科技)
机构由 AI 辅助整理,请以论文原文为准。