arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.19206cs.ARcs.PL

Epic:存储内计算的高效编程范式

Programming In-Storage Computing with Located, Stateful Dataflow

  • UCLA(加州大学洛杉矶分校)

机构由 AI 辅助整理,请以论文原文为准。

Yuyue Wang, Zhenyu Zhang, Glenn Reinman, Huaicheng Li

AI总结:

Epic提出基于NVMe的存储内计算编程范式,通过位置类型、数据流和keep原语统一管理数据驻留与生命周期,实现有状态数据流,平均比最强ISC系统快1.6倍,比主机基线快4.2倍,代码量减少14倍。

AI中文摘要:

存储内计算(ISC)通过在执行计算于计算存储设备(CSD)内部来减少主机与存储之间的数据移动。对于多阶段应用,实现这些益处需要协调工作流中的数据放置、I/O与计算重叠以及设备驻留状态,然而现有接口缺乏对这些决策的统一抽象。我们提出了Epic,一个基于NVMe的ISC栈,通过捕获程序中数据的驻留性和生命周期来提供这种抽象:位置类型声明逻辑驻留性,数据流推导中间值和调用内操作状态的生命周期,以及一个keep原语将选定状态扩展到多次调用之间。这些语义将完整的卸载工作流暴露为一种定位的、有状态的数据流。一个存储感知的编译器转换该工作流,执行移动感知的逻辑映射和融合,并暴露I/O与计算重叠;一个运行时利用执行时信息完成计划,异步地将工作绑定到物理资源并管理设备驻留状态。在12个文件扫描、数据库和机器学习工作负载中,Epic平均比五个先前ISC系统中最强的一个快1.6倍,同时相对于相应主机基线平均实现4.2倍加速,最高达16.1倍,并在我们的实现中将应用侧代码减少多达14倍。

英文摘要:

In-storage computing (ISC) reduces host--storage data movement by executing computation inside computational storage devices (CSDs). For multi-stage applications, realizing these benefits requires coordinating data placement, I/O--compute overlap, and device-resident state across the workflow, yet existing interfaces lack a unified abstraction for these decisions. We present Epic, an NVMe-based ISC stack that provides this abstraction by capturing data residency and lifetime in the program: location types declare logical residency, dataflow derives lifetimes for intermediate values and operation state within an invocation, and a keep primitive extends selected state across invocations. These semantics expose the complete offloaded workflow as a located, stateful dataflow. A storage-aware compiler transforms this workflow, performs movement-aware logical mapping and fusion, and exposes I/O--compute overlap; a runtime completes the plan using execution-time information, asynchronously binding work to physical resources and managing device-resident state. Across 12 file-scanning, database, and machine learning workloads, Epic is 1.6$\times$ faster on average than the strongest of five prior ISC systems, while achieving 4.2$\times$ speedup on average and up to 16.1$\times$ over the corresponding host baselines, and reducing application-side code by up to 14$\times$ in our implementations.

↑