arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

TIDE-MC:面向十亿级GPU矩阵补全的双边插值分解

TIDE-MC: Two-Sided Interpolative Decomposition for Billion-Scale GPU Matrix Completion

Chengying Huan, Yubo Wang, Pinhuan Wang, Lizheng Chen, Jie Zhang, Fangxin Liu, Qing Wang, Ruixuan Liu, Shaonan Ma, Mingxing Zhang, Zhibin Wang, Rong Gu, Guihai Chen, Chen Tian

arXiv 2608.00977首次发表:更新:

AI 中文总结

本文提出基于双边插值分解的GPU框架TIDE-MC,解决现有GPU矩阵补全求解器内存不足或PCIe开销大的问题,可完成超大规模工作负载,实现大幅加速、降内存及低重构误差。

AI 中文摘要

矩阵补全可支撑大规模推荐系统与科学计算,但现有GPU求解器通常假设观测矩阵或其稠密因子能适配设备内存。在实际工作负载下,该假设会导致内存不足(out-of-memory)故障,或在朴素分页时产生严重的PCIe开销。本文提出TIDE-MC,这是一种基于双边插值分解(Two-Sided Interpolative Decomposition, TSID)的有界内存GPU框架。TSID采用采样的模板子矩阵作为锚点来重构完整低秩矩阵,使计算与存储规模随模板和活跃数据块变化,而非随完整矩阵扩展。TIDE-MC通过两个执行阶段实现该方案:第一,无冲突同步引擎利用并行因子分解与分层梯度聚合恢复模板;第二,分块重构流水线将恢复的模板扩展至剩余矩阵,同时将PCIe传输与GPU计算重叠执行。非对称梯度裁剪方案可稳定混合精度Tensor Core执行。在15个基准测试中,TIDE-MC可完成导致现有GPU求解器内存不足的工作负载;与评估的最先进基线相比,其最高实现11647倍加速,峰值内存使用量最高降低8.5倍,重构误差最高降低99.7%。这些结果表明,基于模板的分解与特定阶段的GPU执行可使矩阵补全扩展至超出设备内存容量的规模。

英文摘要

Matrix completion supports large-scale recommendation and scientific computing, yet existing GPU solvers commonly assume that the observed matrix or its dense factors fit in device memory. On real workloads, this assumption leads to out-of-memory failures or severe PCIe overhead under naive paging. We present TIDE-MC, a bounded-memory GPU framework built on Two-Sided Interpolative Decomposition (TSID). TSID uses a sampled template submatrix as an anchor for reconstructing the full low-rank matrix, allowing computation and storage to scale with the template and active data chunks rather than the complete matrix. TIDE-MC realizes this formulation through two execution stages. First, a conflict-free synchronization engine recovers the template using parallel factorization and hierarchical gradient aggregation. Second, a chunked reconstruction pipeline extends the recovered template to the remaining matrix while overlapping PCIe transfers with GPU computation. An asymmetric gradient-clipping scheme stabilizes mixed-precision Tensor Core execution. Across 15 benchmarks, TIDE-MC completes workloads that cause existing GPU solvers to run out of memory. Compared with the evaluated state-of-the-art baselines, it achieves up to 11,647x speedup, reduces peak memory usage by up to 8.5x, and lowers reconstruction error by up to 99.7%. These results show that template-anchored decomposition and stage-specific GPU execution can scale matrix completion beyond device-memory capacity.

Comments15 pages, 12 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑