AI 中文总结
本文针对大图代表性节点采样的现有方法无法适配大规模场景的问题,提出基于最小内积贪心选择规则的列选择性图采样算法,经理论分析与实验验证,该算法适用于大规模图且采样效果良好。
AI 中文摘要
从大图中采样代表性节点是图信号处理与网络分析的基础,但现有方法需访问完整的图拉普拉斯矩阵,在大规模场景下不实用。本文提出一种简单有效的列选择性图采样算法,基于最小内积贪心选择规则,每轮迭代仅访问拉普拉斯矩阵的一小部分随机列,无需特征分解或全局图遍历,适用于无法将完整拉普拉斯矩阵存入内存的大规模图。我们在随机块模型下分析该算法,结果显示当节点度分布均衡时,算法实现与簇大小成比例的采样,所得均值估计对Paley-Wiener空间内带限图信号有效,误差随簇间连接性减弱而降低。在合成与真实数据上的数值实验验证了该方法的有效性。
英文摘要
Sampling representative nodes from large graphs is fundamental to graph signal processing and network analysis, yet existing methods require access to the full graph Laplacian, making them impractical at scale. We propose a simple and effective column-selective graph sampling algorithm based on a minimum inner product greedy selection rule. At each iteration, the algorithm accesses only a small random subset of Laplacian columns, requiring no eigendecomposition or global graph traversal, making it well-suited for large-scale graphs where the full Laplacian cannot be stored in memory. We analyze the algorithm under the stochastic block model and show that, when the degree distribution is balanced across nodes, the algorithm achieves sampling proportional to cluster size, and that the resulting mean estimate is controlled for band-limited graph signals in the Paley-Wiener space, with the error decaying as inter-cluster connectivity weakens. Numerical experiments on both synthetic and real-world data validate the effectiveness of the proposed method.