arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.30497cs.DC

通过稀疏编译统一内存数据分析

Unifying In-Memory Data Analytics through Sparse Compilation

Anand Jayarajan, Gennady Pekhimenko

首次发表
浏览论文内容

中文总结 AI 辅助

Reffine通过基于关系代数和稀疏迭代理论的统一IR及稀疏编译后端,实现跨工作负载的高性能内存数据分析,在TPC-H上比DuckDB快24.9倍,比Umbra快3.2倍,并在流式和图分析上分别比Polars和NetworkX快18.3倍和47.9倍。

中文摘要 AI 辅助

随着现代数据分析工作负载日益异构化和硬件密集型,在多样化应用中实现高效的多核性能仍然是一个开放的挑战。我们提出了Reffine,一个基于编译器的内存分析引擎,可在广泛的数据分析工作负载中提供高性能。Reffine引入了一种新颖的中间表示(IR),基于关系代数和稀疏迭代理论,为数据和计算提供了统一抽象。这种表示支持与工作负载无关的端到端优化,如算子融合和跨多样化分析应用的自动并行化。我们进一步开发了一个稀疏编译器后端,将Reffine IR转换为硬件高效的命令式代码,无需特定领域的实现即可实现高多核性能。在TPC-H基准测试中,Reffine的性能比内存分析数据库DuckDB高出高达24.9倍,比最先进的基于编译的数据库Umbra高出高达3.2倍。Reffine在流式分析和图分析工作负载上,分别比Polars和NetworkX平均加速18.3倍和47.9倍。源代码:此https URL

英文摘要

As modern data analytics workloads become increasingly heterogeneous and hardware-intensive, achieving efficient multi-core performance across diverse applications remains an open challenge. We present Reffine, a compiler-based in-memory analytics engine that delivers high performance across a broad range of data analytics workloads. Reffine introduces a novel intermediate representation (IR), grounded in relational algebra and sparse iteration theory, that provides a unified abstraction for data and computation. This representation enables workload-agnostic, end-to-end optimizations such as operator fusion and automatic parallelization across diverse analytics applications. We further develop a sparse compiler backend that translates Reffine IR into hardware-efficient imperative code, achieving high multi-core performance without domain-specific implementations. On the TPC-H benchmark, Reffine outperforms the in-memory analytical database DuckDB by up to $24.9\times$ and the state-of-the-art compilation-based database Umbra by up to $3.2\times$. Reffine also achieves average speedups of $18.3\times$ and $47.9\times$ over Polars and NetworkX on streaming and graph analytics workloads, respectively. Source code: https://github.com/ampersand-projects/reffine

发表机构

  • University of Toronto(多伦多大学)
  • NVIDIA(英伟达)

机构由 AI 辅助整理,请以论文原文为准。

↑