arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

当词汇变化产生误导时:结合传统指标与基于大语言模型的指标重新思考动态主题模型评估

When Lexical Change Misleads: Rethinking Dynamic Topic Model Evaluation with Traditional and LLM-Based Metrics

Charu Karakkaparambil James

arXiv 2608.13835首次发表:更新:

发表机构

RPTU University Kaiserslautern-Landau(莱茵兰-普法尔茨州立大学凯撒斯劳滕-兰道分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对动态主题模型评估中词汇变化误导的问题,对比传统指标与基于LLM的语义相似性指标,发现后者在CoNTM上与人工判断一致性更高,提出应结合两类指标开展感知词汇变化的评估。

AI 中文摘要

动态主题模型用于捕捉不断演变的词分布,但当词汇发生变化而语义保持不变时,传统的一致性指标可能失效。我们评估了CoNTM和DLDA在《纽约时报》(NYT)、DBLP及arXiv数据集上的120个主题,采用3名人工标注员及低、中、高三个词汇变化类别。传统时间一致性与人工判断的一致性差异极大(相关系数ρ=-0.256至0.614);相比之下,基于大语言模型(LLM)的语义相似性指标对CoNTM在NYT(ρ=0.609)、DBLP(ρ=0.721)及arXiv(ρ=0.502)上与人工语义判断高度一致,但对DLDA的一致性较低。词汇变化分层分析揭示了聚合评估所掩盖的差异,因此我们提倡采用感知词汇变化的评估方式,同时报告传统一致性与基于LLM的语义指标,二者为互补而非可互换的信号。

英文摘要

Dynamic topic models capture evolving word distributions, but traditional coherence metrics may fail when vocabulary changes while semantic meaning persists. We evaluate 120 topics from CoNTM and DLDA across NYT, DBLP, and arXiv, using three human annotators and Low, Medium, and High lexical-change categories. Traditional temporal coherence shows highly variable agreement with human judgments ($ρ$=-0.256 to 0.614). In contrast, LLM-based semantic similarity agrees strongly with human semantic judgments for CoNTM on NYT ($ρ$=0.609), DBLP ($ρ$=0.721), and arXiv ($ρ$=0.502), but is less consistent for DLDA. Lexical-change stratification reveals variation hidden by aggregate evaluation. We therefore advocate lexical-change-aware evaluation, jointly reporting traditional coherence and LLM-based semantic measures as complementary rather than interchangeable signals.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑