arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.08088cs.CLcs.SE

Vectorizer:基于形状引导重写的NumPy程序向量化

Vectorizer: Vectorizing NumPy Programs with Shape-Guided Rewrite

Jingqian Liu, Xiaoyu Liu, Yuepeng Wang

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出Vectorizer工具,通过形状引导的源到源重写将NumPy显式循环程序自动向量化,在150个基准上成功向量化144个,平均加速74.83倍。

中文摘要 AI 辅助

NumPy是一个广泛使用的Python数值科学计算库,以其声明式API和优化实现而闻名。然而,编写高效的NumPy程序(通常需要使用向量化数组操作而非显式Python循环)可能并不简单。对于习惯于命令式数组遍历的程序员来说,这尤其困难,特别是当向量化API调用需要仔细推理形状、广播和高级索引时。本文提出了一种基于重写的方法,用于对具有显式数组数据循环的NumPy程序进行向量化。我们的方法从内到外对循环进行向量化,利用数组形状和数据流分析来指导源到源转换,将循环体替换为向量化语句。遵循一组构造上正确的重写规则,我们的方法始终快速。我们已将该方法实现为一个名为Vectorizer的工具,并在从先前工作和Stack Overflow收集的150个基准上进行了评估。评估显示,Vectorizer直接向量化了150个基准中的142个,在原始基准稍作修改后额外向量化了2个,平均每个重写仅需0.53秒。生成的程序平均比原始的基于循环的实现快74.83倍。

英文摘要

NumPy is a widely used Python library for numerical scientific computing, known for its declarative APIs and its optimized implementations. However, writing efficient NumPy programs, which often entails using vectorized array operations instead of explicit Python loops, may not be straightforward. This can be difficult for programmers who are accustomed to imperative array traversal, especially when vectorized API invocations require careful reasoning about shapes, broadcasting, and advanced indexing. This paper presents a rewrite-based approach for vectorizing Numpy programs with explicit loops over array data. Our approach vectorizes loops from the inside out, using array shapes and dataflow analysis to guide a source-to-source transformation that replaces loop bodies with vectorized statements. Following a set of rewrite rules that are correct by construction, our approach is consistently fast. We have implemented the approach as a tool called Vectorizer and evaluated it on 150 benchmarks collected from prior work and Stack Overflow. The evaluation shows that Vectorizer vectorizes 142 of the 150 benchmarks directly and 2 more after minor changes to the original benchmarks, with only 0.53 seconds on average to rewrite each one. The resulting programs are, on average, 74.83x faster than the original loop-based implementations.

发表机构

  • Simon Fraser University(西蒙菲莎大学)

机构由 AI 辅助整理,请以论文原文为准。

↑