面向跨语料时间分析的动态主题模型
Dynamic Topic Modeling for Cross-Corpus Temporal Analysis
浏览论文内容
中文总结 AI 辅助
针对跨语料时间分析中主题对应不稳定的问题,提出带共享主干和残差适配的D-ETM框架,在三个97年跨度的语料上验证其对齐效果显著优于全微调及独立训练方法。
中文摘要 AI 辅助
动态嵌入主题模型(Dynamic Embedded Topic Models,D-ETM)提供了一种可解释的框架,用于建模时间语义演化,但跨语料比较仍然困难,因为主题通常是独立学习的,仅在训练后才进行对齐,该过程无法保证跨语料和时间的主题对应关系稳定。为解决此问题,我们提出一种D-ETM框架:首先在合并的多语料集合上学习一个通用动态主题空间,我们称之为共享主干(shared backbone);然后在冻结的主干周围引入语料特定的残差适配,无需创建单独的潜在主题空间。该设计保留了用于跨语料比较的共享主题索引,同时允许每个语料在词汇层面进行专业化调整。我们在三个跨越97年的时间结构化语料上评估该框架:美国历史英语语料库(Corpus of Historical American English)、《哈佛商业评论》(Harvard Business Review)和《国际劳工评论》(International Labour Review)。残差适配相对于共享主干提升了语料特定的拟合度,同时保留了同索引跨语料主题轨迹,实现了比从同一主干进行全微调强得多的对齐效果,轨迹检索@1(trajectory Retrieval@1)的结果为97.5±0.7%,而全微调仅为17.9±1.1%,且比采用事后匈牙利匹配(Hungarian matching)的独立训练的对齐效果更强。这些结果表明,将主题对齐纳入模型可支持更稳定的跨语料时间比较,同时保留语料特定的词汇变异。
英文摘要
Dynamic Embedded Topic Models (D-ETM) provide an interpretable framework for modeling temporal semantic evolution, but cross-corpus comparison remains difficult because topics are often learned independently and aligned only after training, a process that does not guarantee stable topic correspondence across corpora and time. To address this problem, we propose a D-ETM framework that first learns a common dynamic topic space over a merged multi-corpus collection, which we call the shared backbone, then introduces corpus-specific residual adaptation around the frozen backbone without creating separate latent topic spaces. This design preserves a shared topic index for cross-corpus comparison while allowing each corpus to specialize lexically. We evaluate the framework on three temporally structured corpora spanning 97 years: the Corpus of Historical American English, Harvard Business Review, and International Labour Review. Residual adaptation improves corpus-specific fit relative to the shared backbone while preserving the same-index cross-corpus topic trajectories, achieving substantially stronger alignment than full fine-tuning from the same backbone, with $97.5 \pm 0.7\%$ versus $17.9 \pm 1.1\%$ trajectory Retrieval@1, as well as stronger alignment than independent training with post-hoc Hungarian matching. These results suggest that incorporating topic alignment into the model can support more stable over-time cross-corpus comparisons while retaining corpus-specific lexical variation.
发表机构
- Columbia University(哥伦比亚大学)
机构由 AI 辅助整理,请以论文原文为准。