发表机构
The University of Tokyo(东京大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出 vSLDA,一种基于句子嵌入的球面主题模型,利用 vMF 分布匹配余弦几何,线性复杂度优于高斯模型,在细粒度主题划分中优于八个基线。
AI 中文摘要
潜在狄利克雷分配(LDA)及其衍生模型仍然是广泛使用的主题模型。LDA 将每篇文档视为词袋,并通过词汇表上的类别分布对每个主题建模,因此它既未利用文档的内部结构,也未利用词语之间的语义相似性。早期工作以两种方式回应了这一局限:许多模型引入了嵌入,一些模型将主题分配给句子而非单词。而两者的结合——一种观察句子嵌入的主题模型——仍鲜有探索。我们提出了 vMF 句子 LDA(vSLDA),它保留了 LDA 的混合结构,将每个句子视为其 L2 归一化嵌入,并通过 von Mises-Fisher(vMF)分布对每个主题建模,这与句子嵌入的余弦几何相匹配。其每个主题的参数数量和每次迭代的成本与嵌入维度呈线性关系,而现有基于句子嵌入的模型中的全协方差高斯分布则是二次关系。我们在该局限预期影响最大的场景中评估 vSLDA,即共享大量词汇的主题之间:细分一个集合的单一主题的子主题,以及当语料库被划分为大量主题时产生的窄主题。在两个语料库上,当从文档-主题分布中分类每个粗类别内的细类别时,vSLDA 在八个基线中取得了最佳平均排名。在整个语料库上,随着主题数量的增加,其优势显现或扩大。通过每个句子的主题后验对其词频进行加权,可得到与 LDA 相同形式的期望主题-词计数,因此标准的主题连贯性和多样性度量适用于将主题分配给句子的模型。在它们的乘积——主题质量上,vSLDA 在两个语料库的大多数类别内条件下均领先。
英文摘要
Latent Dirichlet Allocation (LDA) and models derived from it remain widely used topic models. LDA observes each document as a bag-of-words and models each topic by a categorical distribution over the vocabulary, so that it uses neither the internal structure of the document nor the similarity in meaning between words. Earlier work has responded to this limitation in two ways: many models have introduced embeddings, and some have assigned topics to sentences rather than to words. Their combination, a topic model that observes sentence embeddings, remains little explored. We propose vMF Sentence LDA (vSLDA), which keeps the admixture structure of LDA, observes each sentence as its L2-normalized embedding and models each topic by a von Mises-Fisher (vMF) distribution, which matches the cosine geometry of sentence embeddings. Its per-topic parameter count and per-iteration cost are linear in the embedding dimension, versus quadratic for the full-covariance Gaussian distribution in the existing model over sentence embeddings. We evaluate vSLDA where the limitation is expected to matter most, among topics that share much of their vocabulary: the topics that subdivide the one subject of a collection, and the narrow topics that result when a corpus is divided into a large number of topics. On two corpora, vSLDA attains the best mean rank against eight baselines when the fine categories within each coarse category are classified from the document-topic distributions. On the whole corpus, its advantage appears or widens as the number of topics grows. Weighting the word frequencies of each sentence by its topic posterior yields expected topic-word counts of the same form as those of LDA, so that the standard topic coherence and diversity measures apply to models that assign topics to sentences. On their product, topic quality, vSLDA leads in most within-category conditions of both corpora.