arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

从离群点预测新兴主题:嵌入空间中弱信号的前瞻性研究

Predicting Emerging Topics from Outliers: A Prospective Study of Weak Signals in Embedding Space

Evangelia Zve, Gauvain Bourgne, Jean-Gabriel Ganascia

arXiv 2609.29183首次发表:更新:

发表机构

LIP6, Sorbonne Université, CNRS(索邦大学计算机科学实验室LIP6,法国国家科学研究中心)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究利用嵌入空间中离群点的几何特征,前瞻性预测新兴主题的弱信号,在两个法语新闻语料库上验证了方法的有效性,F1值最高超过0.90。

AI 中文摘要

一些最初被基于嵌入的主题模型归类为噪声的文档,后来成为新兴主题的创始成员。然而,在发表时,它们在嵌入空间中表现为分散的点,且在没有后见之明的情况下难以与普通噪声区分。我们研究是否可以利用文档首次出现时仅有的信息,前瞻性地预测这类预期性离群点。我们从离群点的后续轨迹中推导标签,区分那些预示新主题的离群点与那些强化现有主题或保持孤立的离群点,并通过多个嵌入模型之间的一致性来估计标签置信度。在两个法语新闻语料库上,预期性离群点在发表时被证明是可预测的。在交叉验证下,F1值从整个合格人群的约0.77上升到高一致性子集上的0.90以上,并且在严格按时间顺序的评估下保持在0.76-0.80。预测性能主要由捕捉每个离群点在嵌入空间中位置的几何特征驱动。

英文摘要

Some documents that embedding-based topic models initially classify as noise later become founding members of emerging topics. At publication time, however, they appear as scattered points in embedding space and are difficult to distinguish from ordinary noise without the benefit of hindsight. We study whether such anticipatory outliers can be predicted prospectively, using only information available when a document first appears. We derive labels from the subsequent trajectories of outlier documents, distinguishing those that anticipate new topics from those that reinforce existing topics or remain isolated, and estimate label confidence through agreement across multiple embedding models. On two French news corpora, anticipatory outliers prove predictable at publication time. Under cross-validation, $F_1$ rises from about 0.77 over the full eligible population to above 0.90 on high-consensus subsets, and remains at 0.76-0.80 under a strictly chronological evaluation. Predictive performance is driven mainly by geometric features capturing each outlier's position in embedding space.

CommentsAccepted to Findings of AACL-IJCNLP 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑