NDS:高性能近数据线程的无程序员卸载
NDS: Programmer-Free Offload of High-Performance Near Data Strands
- University of Utah(犹他大学)
- Intel(英特尔)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文提出NDS框架,通过硬件中心方法自动识别并卸载规则循环至近数据处理,无需程序员或软件栈修改,实现单核3.14倍和8核1.82倍的几何平均加速。
AI中文摘要:
近数据处理(NDP)有潜力通过缓解数据移动瓶颈来显著提升系统性能和能效。然而,大多数NDP方案对软件栈、数据布局和/或底层硬件提出了较高要求。为了扩大NDP的采用范围,本工作聚焦于一种模块化的硬件中心方法,该方法自动化上述步骤,并能针对未修改的二进制文件。本文提出了近数据线程(NDS),一个无需程序员和软件栈参与即可自动编排任务和数据的框架。作为实现这一宏大目标的第一步,本工作聚焦于满足特定条件的规则循环。该框架识别潜在可卸载的循环,将其分解为并行执行线程,创建高性能指令调度,执行所需的数据编组和编排,并启动近数据执行。我们证明,NDS相对于单核主机基线实现了3.14倍的几何平均加速,相对于8核主机并行变体实现了1.82倍的几何平均加速,所有这些均无需对软件栈进行任何修改。随着重复循环调用摊销卸载成本,这些增益进一步增长,表明遗留应用可以利用透明的硬件方法从NDP中获益。
英文摘要:
Near Data Processing (NDP) has the potential to significantly improve system performance and energy by alleviating data movement bottlenecks. However, most NDP proposals pose heavy requirements for the software stack, data layout, and/or the underlying hardware. To broaden NDP adoption, this work focuses on a modular hardware-centric approach that automates the above steps and can target unmodified binaries. This paper presents Near Data Strands (NDS), a framework that orchestrates tasks and data automatically without involving the programmer and the software stack. As a first step towards this ambitious goal, this work focuses on regular loops that meet specific criteria. The framework identifies potential offloadable loops, decomposes them into parallel execution strands, creates a high performance instruction schedule, performs the required data marshaling and orchestration, and initiates near-data execution. We demonstrate that NDS achieves a 3.14x geomean speedup over a single-core host baseline and a 1.82x geomean speedup over 8-core host-parallel variants, all without any modifications to the software stack. These gains grow as repeated loop invocations amortize offload costs, showing that legacy applications can leverage a transparent hardware approach to extract benefits from NDP.