Reddit语料库中评论级别的主题漂移分析
Comment-level Topic Drift Analysis in the Reddit Corpus
浏览论文内容
中文总结 AI 辅助
该研究将嵌入式动态主题建模技术应用于Reddit语料库,分析127亿条2006-2022年的评论,提出嵌入空间内语义漂移分析方法,发现争议性主题漂移显著而文体类主题相对稳定。
中文摘要 AI 辅助
我们提出了一种基于嵌入的动态主题建模技术的新应用,用于检测和量化大规模语料库中评论级别的主题漂移。通过利用预训练语言模型为短文本生成上下文语义嵌入,我们分析了2006年至2022年间的127亿条Reddit评论。我们对这些嵌入采用无监督方法,识别出随时间动态演变的主题簇。我们的主要贡献是一种在嵌入空间本身分析语义漂移和话语演变的方法。我们还展示了对现有方法的修改,以实现大规模分析,并提出并演示了一种零模型比较测试来过滤虚假动态。关键发现表明,政治和社会争议性主题在嵌入空间中表现出显著的定向漂移,主题间距离随时间发生系统性变化,超出了零模型所能解释的范围,而音乐和体育等领域则保持相对稳定。
英文摘要
We present a novel application of embedding-based dynamic topic modeling techniques to detect and quantify topic drift at the comment level in a massive corpus. By leveraging pretrained language models to generate contextualized semantic embeddings for short text, we analyzed 12.7 billion Reddit comments spanning 2006 to 2022. Using unsupervised methods on these embeddings, we identify dynamically evolving topic clusters over time. Our primary contribution is a methodology for analysis of semantic drift and discourse evolution in the embedding space itself. We also demonstrate modifications to existing methods that enable this analysis at scale, and we propose and demonstrate a null model comparison test to filter spurious dynamics. Key findings suggest that politically and socially contentious topics exhibit significant directional drift in embedding space, with inter-topic distances changing systematically over time beyond what the null model can explain, whereas domains such as music and sports remain comparatively stable.
发表机构
- William & Mary(威廉玛丽学院)
机构由 AI 辅助整理,请以论文原文为准。