发表机构
Clemson University(克莱姆森大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出两种随机化SVD近似方法用于词-文档矩阵的谱共聚类,通过随机投影和随机采样降低计算成本,实验表明方法选择应依据矩阵稀疏性。
AI 中文摘要
谱共聚类是发现词-文档矩阵中潜在结构的有用工具,但其对奇异值分解(SVD)的依赖可能使标准公式在高维数据上计算代价高昂。本文针对文档簇和词簇数量可能不同的二分文本数据,提出了两种归一化谱共聚类的随机化近似方法。第一种方法通过随机投影使用随机化SVD,第二种方法则将部分SVD与逐元素随机采样相结合。在真实世界和合成数据集上,两种方法相对于完整SVD基线均减少了运行时间,但其行为取决于矩阵的稀疏性。在所测试的设置中,随机投影方法是更可靠的近似,而基于采样的方法在较稠密的矩阵上最为有用,在已经稀疏的文本数据上提供的益处有限。这些结果表明,应根据数据的底层结构来选择用于谱共聚类的随机化近似方法。
英文摘要
Spectral co-clustering is a useful tool for discovering latent structure in word-document matrices, but its reliance on singular value decomposition (SVD) can make standard formulations expensive on high-dimensional data. This paper presents two randomized approximations for normalized spectral co-clustering of bipartite text data when the numbers of document and word clusters may differ. The first method uses randomized SVD through random projection, while the second combines partial SVD with element-wise random sampling. Across real-world and synthetic datasets, both methods reduce runtime relative to the full-SVD baseline, but their behavior depends on matrix sparsity. The random projection method is the more reliable approximation across the tested settings, whereas the sampling-based method is most useful on denser matrices and provides limited benefit on already sparse text data. These results show that randomized approximations for spectral co-clustering should be selected according to the underlying structure of the data.
CommentsAccepted to 2026 IEEE International Conference on Machine Learning and Applications (ICMLA)