arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.06328cs.PL

SparseConflicts:处理稀疏张量收缩中的冲突数据布局

SparseConflicts: Handling Conflicting Data Layouts in Sparse Tensor Contractions

Adhitha Dias, Kirshanthan Sundararajah, Artem Pelenitsyn, Milind Kulkarni

首次发表
浏览论文内容

中文总结 AI 辅助

针对稀疏张量收缩中数据布局冲突导致的显式转置开销,本文扩展TACO编译器中间表示,生成避免转置的融合循环,实现高达2倍加速。

中文摘要 AI 辅助

优化稀疏张量计算具有挑战性,因为压缩存储格式的使用导致了非仿射循环嵌套和庞大而复杂的调度空间。给定调度的性能对输入张量的稀疏模式敏感,使得难以找到单一最优解。当同一张量收缩中的输入张量相对于迭代顺序具有冲突的数据布局时,需要代价高昂(在时间和内存方面)的布局转换,例如转置。一种有前景但尚未充分探索的替代方案是生成避免显式转置的调度,但现有编译器尚未系统性地支持这一点。本文提出了一种新的代码生成策略,该策略推广了TACO稀疏张量编译器的中间表示,以生成单个循环嵌套,从而在张量具有冲突的数据布局迭代顺序时规避显式转置。我们扩展了TACO的迭代图,以表达一种基于搜索的策略,用于定位具有冲突布局的张量中的元素,并引入了新的中间表示节点,以将这些调度降级为高效代码。这使得系统性地生成不需要显式转置的循环成为可能,从而避免了物化临时张量的开销。我们使用真实世界和合成数据集对一组稀疏张量收缩评估了我们的方法。结果表明,对于数据布局未对齐的计算,我们的融合方法在某些稀疏模式下比传统显式创建转置临时张量的方法实现了高达2倍的加速。我们还提供了关于这种新调度策略何时可能有益的指导方针。

英文摘要

Optimizing sparse tensor computations is challenging due to the use of compressed storage formats, which leads to non-affine loop nests and a vast, complex schedule space. The performance of a given schedule is sensitive to the sparsity pattern of the input tensors, making it difficult to find a single optimal solution. When input tensors in the same tensor contraction have conflicting data layouts in relation to the iteration order, it requires costly-both in time and memory-layout transformation, such as transposition. A promising but under-explored alternative is to generate a schedule that avoids explicit transposition, but this has not been systematically supported in existing compilers. This paper presents a new code generation strategy that generalizes the intermediate representation of the TACO sparse tensor compiler to generate a single loop nest, circumventing explicit transposition of tensors when the tensors have conflicting data layout iteration orders. We extend TACOś iteration graph to express a search-based strategy for locating elements in tensors with conflicting layouts, and we introduce new intermediate representation nodes to lower these schedules to efficient code. This enables the systematic generation of loops that do not require explicit transposition, thus avoiding the overhead of materializing temporary tensors. We evaluate our approach on a set of sparse tensor contractions using both real-world and synthetic datasets. Our results demonstrate that for computations with misaligned data layouts, our fused approach achieves up to 2x speedup for some sparsity patterns over the traditional approach of explicitly creating a transposed temporary. We also provide guidelines for when this new scheduling strategy is likely to be beneficial.

发表机构

  • Purdue University(普渡大学)
  • Virginia Tech(弗吉尼亚理工大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑