发表机构
IBM Research; MIT(IBM研究院; 麻省理工学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出一种更精简的Transformer,以更小的嵌入大小执行k-means聚类算法,并训练其学习聚类任务,理论刻画和实证验证了影响收敛与泛化的因素,同时探究了其通用聚类能力的成败情形。
AI 中文摘要
Transformer具有上下文学习能力,其中一些已知的学习算法可以通过模型的前向传播来执行。近期工作表明,Transformer可以精确执行Lloyd算法,用于对d维空间中的n个点进行k-means聚类,其嵌入大小为d_emb = d + k(因此,需要大小为(d+k)^2的注意力投影矩阵)。在本工作中,我们基于此结果进行了以下扩展:首先,我们提出了一个表达能力相同但更小的Transformer,它以嵌入大小d_emb = (d + ⌈log2 k⌉)执行Lloyd算法。接下来,我们训练这些Transformer以在给定聚类任务分布的情况下学习聚类算法,并从理论上刻画和实证验证了影响基于随机梯度的学习算法收敛性和分布内泛化的因素。最后,我们探究了这些学习算法(以Transformer形式)的通用聚类能力,并试图理解其成功与失败的情形。
英文摘要
Transformers have in-context learning capabilities, where some known learning algorithms can be executed in the forward pass through the model. Recent work shows that transformers can exactly perform Lloyd's algorithm for $k$-means clustering with $n$ points in $d$ dimensions with an embedding size $d_{\textsf{emb}} = d+k$ (thus, requiring attention projection matrices of size $(d+k)^2$). In this work, we build upon this result in the following ways: First, we present an equally expressive but smaller transformer that executes Lloyd's algorithm with embedding size $d_{\textsf{emb}} = (d + \lceil \log_2 k \rceil)$. Next, we train these transformers to learn the clustering algorithms given a distribution of clustering tasks, and theoretically characterize and empirically validate the factors affecting the convergence and in-distribution generalization of learning algorithms based on stochastic gradients. Finally, we probe the general clustering abilities of these learned algorithms (in the form of transformers), and try to understand situations where they succeed and fail.
CommentsNeurIPS 2026 accepted paper