arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

重尾文档长度下的高效主题模型估计

Efficient Topic Model Estimation under Heavy-Tailed Document Lengths

Daniel Cirkovic, Tiandong Wang

arXiv 2607.24641首次发表:更新:

AI 中文总结

研究重尾文档长度下的主题模型估计,通过证明LDA模型能适应幂律单词频率,利用文档长度分布正则变化时的单词频率幂律层次结构,开发高效张量分解算法,应用于20个新闻组语料库,展现对数据预处理选择的鲁棒性,推动机器学习方法用于多元极值研究。

AI 中文摘要

早期对自然语言统计特性的研究发现,单词出现频率呈幂律分布,这与齐普夫定律相关。但文本这一特性在统计推断中很少被利用。本文证明潜在狄利克雷分配(LDA)模型可适应幂律单词频率。当文档长度分布呈正则变化时,单词频率分布在文档和主题间呈现幂律层次结构。利用此发现开发了通过归一化极端单词频率矩估计主题矩阵的高效张量分解算法。将算法应用于20个新闻组语料库表明,极值方法对数据预处理中的某些选择具有鲁棒性。这项工作推动了机器学习方法在多元极值研究中的应用。

英文摘要

Early inquiries into the statistical properties of natural language found that words tend to occur with power-law frequencies. This observation, closely associated with Zipf's law, has spurred many investigations into why this power-law pattern emerges with such regularity. Rarely, however, has this property of text been leveraged in statistical inference. In this paper, we demonstrate that the Latent Dirichlet Allocation (LDA) model can accommodate power-law word frequencies. In particular, when the document length distribution is regularly varying, the word frequency distribution admits a hierarchy of power laws across documents and topics. We further leverage this finding to develop an efficient tensor decomposition algorithm for estimating the topic matrix via the moments of normalized extreme word frequencies. Applying our algorithm to the twenty newsgroups corpus reveals that the extreme-value methodology exhibits robustness to certain choices made in the pre-processing of the data. This work furthers the recent interest in adapting machine learning methods to the study of multivariate extremes.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑