发表机构
Virginia Tech(弗吉尼亚理工大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
Splyce是MLIR中的自动向量化框架,通过双路径执行和选择性谓词消除稀疏协同迭代中的分支依赖,在合成和真实数据集上实现1.96至2.86倍加速。
AI 中文摘要
稀疏张量收缩的性能瓶颈在于稀疏-稀疏协同迭代循环,这类循环难以进行标准循环向量化。我们提出了Splyce,一个基于MLIR的自动向量化框架,通过双路径执行模型克服了这一挑战。通过选择性谓词将坐标交集与指针管理解耦,Splyce作为副作用固有地消除了数据依赖分支,使现代超标量引擎能够最大化指令级并行并隐藏内存延迟。除了简单的分支消除外,我们的变换还暴露了可并发执行的独立计算,提高了功能单元利用率,而这一利用率原本会受到顺序依赖的约束。在基础稀疏张量内核上的评估表明,在合成输入上性能提升范围为1.96倍至2.86倍,并且在来自SuiteSparse集合的绝大多数不规则真实世界数据集上持续保持一致的加速。最终,Splyce证明了通过将不可预测的控制流转换为可预测的数据流,编译器驱动的推测能够有效调和压缩存储的内存效率与现代超标量架构的执行单元吞吐量。
英文摘要
Sparse tensor contractions are bottlenecked by sparse-sparse coiteration loops that resist standard loop vectorization. We present Splyce, an auto-vectorization framework in MLIR that overcomes this through a dual-path execution model. By decoupling coordinate intersection from pointer management via selective predication, Splyce inherently eliminates data-dependent branches as a side effect, allowing modern superscalar engines to maximize instruction-level parallelism and hide memory latency. Beyond simple branch elimination, our transformation exposes independent computation that can be executed concurrently, increasing functional-unit utilization that would otherwise be constrained by sequential dependencies. Evaluation across foundational sparse tensor kernels demonstrates performance ranging from 1.96X to 2.86X on synthetic inputs, with consistent speedups sustained across a vast majority of irregular real-world datasets from the SuiteSparse collection. Ultimately, Splyce demonstrates that by converting unpredictable control-flow into a predictable data stream, compiler-driven speculation can effectively reconcile the memory efficiency of compressed storage with the execution-unit throughput of modern superscalar architectures.