发表机构
MEF University(MEF大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本工作提出强化学习引导的图变换框架,将SpTRSV优化建模为顺序决策问题,学习矩阵依赖的变换策略,实现级别数平均减少23%、级别成本变异系数平均减少29%,并支持策略迁移。
AI 中文摘要
稀疏三角求解(SpTRSV)是众多科学与工程应用中的基础内核。然而,稀疏三角矩阵中固有的数据依赖显著限制了可用的并行性,并使高效的工作负载分配变得具有挑战性。最近的图变换技术通过修改输入矩阵的依赖图来改善并行执行,从而解决了这些限制。然而,现有的图变换策略依赖于手工设计的启发式方法,这使得它们的发展和适应不同优化目标变得困难。本工作提出了一种用于SpTRSV的强化学习引导的图变换框架,其中图变换被表述为一个顺序决策问题,并且一个RL智能体学习依赖于矩阵的变换策略。在真实世界稀疏矩阵上的实验结果表明,级别数最多减少94%,级别成本的变异系数最多减少80%,而在最高情况下仅修改了1.50%的行。平均而言,RL引导的图变换实现了级别数减少23%,级别成本的变异系数减少29%,同时仅重写了0.82%的矩阵行。尽管启发式策略通常实现更激进的级别减少(在31%到46%之间),但基于RL的方法在级别成本的变异系数上实现了最大的平均减少,展示了其平衡竞争性图变换目标的能力。结果进一步表明,通过课程学习和微调,学习到的策略可以迁移到先前未见过的矩阵上,而零样本实验则提供了关于图变换策略在不同稀疏模式间泛化局限性的见解。
英文摘要
Sparse triangular solve (SpTRSV) is a fundamental kernel in numerous scientific and engineering applications. However, the data dependencies inherent in sparse triangular matrices significantly limit the available parallelism and make efficient workload distribution challenging. Recent graph transformation techniques address these limitations by modifying the dependency graph of the input matrix to improve parallel execution. Existing graph transformation strategies, however, rely on manually designed heuristics, making their development and adaptation to different optimization objectives challenging. This work proposes a reinforcement learning-guided graph transformation framework for SpTRSV, in which graph transformation is formulated as a sequential decision-making problem and an RL agent learns matrix-dependent transformation policies. Experimental results on real-world sparse matrices demonstrate level reductions of up to 94% and reductions of up to 80% in the coefficient of variation of level costs, while modifying only 1.50% of the rows in the highest case. On average, the RL- guided graph transformation achieves a 23% reduction in the number of levels and a 29% reduction in the coefficient of variation of level costs while rewriting only 0.82% of the matrix rows. Although the heuristic strategies generally achieve more aggressive level reduction(between 31% and 46%), the RL-based approach achieves the largest average reduction in the coefficient of variation of level costs, demonstrating its ability to balance competing graph transformation objectives. The results further show that the learned policies can be transferred to previously unseen matrices through curriculum learning and fine-tuning, while zero-shot experiments provide insights into the limitations of generalizing graph transformation policies across different sparsity patterns.
Comments33 pages, 3 figures, 7 tables. Submitted to The Journal of Supercomputing and currently under review