AI 中文总结
研究在GPU内存受限场景下语法压缩矩阵右乘问题,提出基于特定有向无环图及分层语法的方法,实现流式评估,在基因型矩阵和软件遗产图上有空间优势、时间性能良好且能量较低。
AI 中文摘要
语法压缩矩阵(mm-repair族)将矩阵的非零结构存储为RePair直线程序(SLP),支持与压缩大小成比例的时间和空间内的矩阵向量乘积。我们针对在GPU上起决定性作用的情况:当未压缩矩阵超过设备内存时,因此占用空间(而非浮点吞吐量)是约束条件。我们的SLP是出度为2的有向无环图(DAG),右积y = Mx是一次自底向上的扫描(从叶到根):无冲突收集。通过直通完成使语法正确分层(每个非终结子节点比其父节点低一级),插入恒等节点向上传递值直到被使用。这产生了一种流式评估,其中每一层仅读取下一层并写入下一层,因此活动集适合两个交替的只读/只写缓冲区,而不是随整个语法扩展;每层宽度等于活动集。在基因型矩阵上,多基因得分恰好是右积y = Gβ,CUDA实现显示出明显的空间优势:设备占用空间比物化的cuSPARSE CSR基线小4到8倍,单向量时间与cuSPARSE相差不大,且能量始终较低。由于扫描仅需要一个关联组合,相同的引擎和调度通过交换小的叶/组合/发射策略来评估语法上的任何幺半群同态;相同的可达性扫描然后扩展到具有十亿条边的软件遗产图(261 TB密集且不可物化,比CSR序列化小21倍),内存论据成立。我们将此构建为一个算法工程案例研究:测量结构指标(深度、活动集宽度、完成成本),这些是与架构无关的语法属性,而时间和能量在单板上进行分析。
英文摘要
Grammar-compressed matrices (the mm-repair family) store a matrix's non-zero structure as a RePair straight-line program (SLP), supporting matrix-vector products in time and space proportional to the compressed size. We target the regime where this is decisive on a GPU: when the uncompressed matrix exceeds device memory, so footprint (not floating-point throughput) is the binding constraint. Our SLP is a directed acyclic graph (DAG) of out-degree 2, and the right product $y=Mx$ is a single bottom-up sweep (leaves to roots): a conflict-free gather. We make the grammar properly layered (every nonterminal child one level below its parent) via pass-through completion, which inserts identity nodes to carry values upward until consumed. This yields a streaming evaluation in which each level reads only the level below and writes the next, so the live set fits in two alternating read-only/write-only buffers instead of scaling with the whole grammar; the per-level width equals the live set. On genotype matrices, where a polygenic score is exactly the right product $y=Gβ$, a CUDA implementation shows a clear space advantage: a device footprint 4 to 8 times smaller than a materialized cuSPARSE CSR baseline, single-vector times within a small factor of cuSPARSE, and consistently lower energy. Because the sweep needs only an associative combine, the same engine and schedule evaluate any monoid homomorphism over the grammar by swapping a small leaf/combine/emit policy; the same reachability sweep then scales to the billion-edge Software Heritage graph ($261$ TB dense and unmaterializable, $21\times$ smaller serialized than CSR), where the memory argument holds. We frame this as an algorithm-engineering case study: structural metrics (depth, live-set width, completion cost) are measured, architecture-independent grammar properties, whereas time and energy are profiled on a single board.
Commentstwo columns, 15 pages, 5 figures, 8 tables