arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

用于高效标签传播的代数多重网格加速

Algebraic Multigrid Acceleration for Efficient Label Spreading

Antonia van Betteray, Jonathan Klees, Miriam Schäfers, Matthias Rottmann

arXiv 2608.26309首次发表:更新:

发表机构

Osnabrück University; Ruhr University Bochum(奥斯纳布吕克大学; 波鸿鲁尔大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对标签传播在大规模高维数据集上的计算与内存局限,提出AMELS框架,通过代数多重网格求解器替代随机游走迭代,大幅减少运行时间且对超参数更鲁棒,可高效应用于大规模图像数据集。

AI 中文摘要

现代机器学习模型依赖大量标注数据,但大规模数据集的人工标注成本高且耗时。标签传播是一种半监督学习技术,通过将少量标注样本的信息传播到大量未标注数据来应对这一挑战。尽管该技术有效,但它在大规模高维数据集上的应用受限于计算成本和内存约束。为解决这些局限,我们提出了用于高效标签传播的代数多重网格加速(AMELS),这是一种高效的标签传播框架,通过快速构建邻接图并融入代数多重网格求解器来提升可扩展性。代数多重网格求解器是一种迭代求解器,可替代标签传播中通常执行的普通随机游走迭代。由于代数多重网格求解器的多级特性,AMELS可在单个多重网格循环中跨任意规模的图传播给定标注信息。我们证明,与现有实现相比,AMELS实现了显著的运行时间减少,同时在运行时间和分类精度方面对超参数选择更具鲁棒性。因此,我们的框架可在大规模图像数据集上实现高效标签传播,且即使仅使用少量标注样本也能生成准确标注。

英文摘要

Modern machine learning models rely on large amounts of labeled data. However, manual annotation of large-scale datasets is expensive and time-consuming. Label spreading is a semi-supervised learning technique that addresses this challenge by propagating information from a few labeled examples to a larger pool of unlabeled data. Despite its effectiveness, its application to large-scale, high-dimensional datasets is limited by computational costs and memory constraints. To address these limitations, we propose Algebraic Multigrid Acceleration for Efficient Label Spreading (AMELS), an efficient label spreading framework that improves scalability by fast construction of neighborhood graphs and the incorporation of algebraic multigrid solvers. The latter is an iterative solver that replaces the ordinary random walk iteration typically performed in label spreading. Due to the multilevel nature of algebraic multigrid solvers, AMELS spreads given label information across a graph of any size in a single multigrid cycle. We demonstrate that AMELS achieves significant runtime reductions compared to existing implementations while also being more robust to hyperparameter choices in terms of both runtime and classification accuracy. Our framework therefore enables efficient label spreading on large-scale image datasets and produces accurate labels even when only a few labeled samples are available.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑