发表机构
State University of São Paulo (UNESP)(圣保罗州立大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出免训练的MARETopic框架,通过基于秩的原型选择在流形上发现主题,无需梯度更新即可在多个基准上提升纯度和NMI,并加速运行。
AI 中文摘要
近年来的主题模型利用预训练嵌入,但神经架构产生的潜在表示缺乏对特定文本的锚定,而基于聚类的流程仅事后指派代表性文档,依赖于高维空间中因枢纽性和各向异性而失真的绝对距离。我们提出MARETopic,一种免训练框架,将主题发现视为基于秩的原型选择。在将嵌入投影到低维流形后,MARETopic构建编码序数邻域结构的排序列表。一种贪心算法精确选择K个示例文档,即真实语料文本,其邻域覆盖整个语料库。两种变体共享此标准。MARETopic$_\text{Corr}$使用查询性能预测器和秩相关度量对候选进行评分,在两个类别最多的基准上领先Purity和NMI,优于神经和基于聚类的主题模型。MARETopic$_\text{Diff}$使用基于秩的扩散矩阵对候选评分,无需任一度量,且运行速度快1.7至1.9倍。无需任何梯度更新,MARETopic在三个数据集中的两个上引领主题连贯性。一种新颖的主题间最大边际相关性步骤以较小的连贯性代价提高了词汇多样性。我们的代码可从此https URL获取。
英文摘要
Recent topic models leverage pretrained embeddings, but neural architectures produce latent representations without grounding in specific texts, and clustering-based pipelines assign representative documents only post hoc, relying on absolute distances distorted by hubness and anisotropy in high-dimensional spaces. We introduce MARETopic, a training-free framework that casts topic discovery as rank-based prototype selection. After projecting embeddings onto a low-dimensional manifold, MARETopic builds ranked lists encoding ordinal neighborhood structure. A greedy algorithm selects exactly K exemplar documents, real corpus texts, whose neighborhoods cover the corpus. Two variants share this criterion. MARETopic$_\text{Corr}$ scores candidates with a query performance predictor and a rank correlation measure, leading Purity and NMI on the two benchmarks with the most categories, ahead of both neural and clustering-based topic models. MARETopic$_\text{Diff}$ scores them with a rank-based diffusion matrix, needs neither measure, and runs 1.7 to 1.9 times faster. Without a single gradient update, MARETopic leads topic coherence on two of three datasets. A novel inter-topic Maximal Marginal Relevance step raises vocabulary diversity at little cost in coherence. Our code is available at https://github.com/thcastilho/maretopic.
CommentsAccepted at the Main Conference of 2026 Conference on Empirical Methods in Natural Language Processing (EMNLP 2026)