发表机构
KTH Royal Institute of Technology; University of Massachusetts Amherst(KTH皇家理工学院; 马萨诸塞大学阿默斯特分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对图数据中大量相同系数矩阵的线性系统求解,提出CAST方法,用加权随机生成树替换Schur补团,实现无偏近似Cholesky预条件,并引入CAST-ρ控制采样变异性。
AI 中文摘要
扩散估计、排序、半监督学习和网络优化等图数据工作负载通常需要求解许多具有相同系数矩阵的拉普拉斯或对称对角占优M矩阵(SDDM)系统。近似Cholesky预条件子一次消除一个顶点,并存储所得的稀疏近似分解,即因子,其构建成本在这些求解中被摊销。但是,消除一个顶点(即主元)会在其d个活动邻居之间产生一个稠密的Schur补团。我们引入了CAST(规范近似Schur树),它用直接从该团中采样的加权随机生成树替换该团。每个实现都是连通的,且恰好包含d-1条边,而通过将每条选中边的权重乘以其包含在树中的概率的倒数进行重新加权,使得更新无偏。该分布与主元邻居的排序无关,我们证明其杠杆分数边际最小化了无偏逆边际单树估计器中最大的归一化重新加权边贡献。我们还引入了CAST-ρ,它将每个主元邻居替换为ρ个副本,每个副本携带该邻居关联权重的1/ρ份额,在扩展团上采样加权随机生成树,并将副本收缩回原始邻域。所得更新保持无偏且连通,可以在O(ρd)时间内精确采样,并满足归一化局部Schur误差二阶矩的1/ρ界限。因此,增加ρ会降低认证的局部采样变异性,但可能增加构建成本和下游填充。经验上,我们观察到CAST-1是更快的默认选择,而CAST-2在其额外边贡献成本较低时更可取。
英文摘要
Graph-data workloads such as diffusion estimation, ranking, semi-supervised learning, and network optimization often solve many Laplacian or symmetric diagonally dominant M-matrix (SDDM) systems with the same coefficient matrix. Approximate Cholesky preconditioners eliminate vertices one at a time and store the resulting sparse approximate factorization, the \emph{factor}, whose construction cost is amortized across these solves. But eliminating a vertex, the \emph{pivot}, creates a dense Schur-complement clique among its $d$ active neighbors. We introduce CAST (Canonical Approximate Schur Tree), which replaces this clique with a weighted random spanning tree sampled directly from it. Every realization is connected and contains exactly d-1 edges, while reweighting each selected edge by the reciprocal of its tree-inclusion probability makes the update unbiased. The distribution is independent of the ordering of the pivot neighbors, and we prove that its leverage-score marginals minimize the largest normalized reweighted-edge contribution among unbiased inverse-marginal one-tree estimators. We also introduce CAST-$ρ$, which replaces each pivot neighbor with $ρ$ copies, each carrying a $1/ρ$ share of that neighbor's incident weight, samples a weighted random spanning tree on the expanded clique, and contracts the copies back to the original neighborhood. The resulting update remains unbiased and connected, can be sampled exactly in $O(ρd)$ time, and satisfies a $1/ρ$ bound on the second moment of the normalized local Schur error. Increasing $ρ$ therefore reduces certified local sampling variability, but may increase construction cost and downstream fill. Empirically, we observe that CAST-1 is the faster default, whereas CAST-2 is preferable when its additional edge contributions remain inexpensive.