arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.16269cs.CLcs.LG

基于上下文词级语义图表示的领域无关神经主题模型

Domain-Agnostic Neural Topic Modeling with Contextual Token-Level Semantic Graph Representation

Seung-Won Seo, Won Ik Cho, Yongmin Yoo

首次发表
浏览论文内容

中文总结 AI 辅助

该研究提出领域无关框架DARTopic,通过构建词级语义图并联合训练GNN编码器与主题推理模块,在多领域基准测试中优于基线且效率更优。

中文摘要 AI 辅助

近期,结合预训练语言模型(PLMs)的神经主题模型借助通用领域预训练取得了优异性能,但在专业语料库上的主题可解释性往往会下降。这一限制主要源于嵌入空间的几何特性:预训练期间未出现的领域特定术语会坍缩到难以区分的区域,而领域特定的再训练、词级图增强或参数高效微调均无法在不继承底层编码器容量上限的前提下重构该空间。我们的核心洞见是:在词级PLM嵌入上运行的可学习图层能够获取冻结编码器所缺失的语料库特定语义结构,因为词级图保留了词级表示所丢弃的文档局部上下文,且与主题目标的联合优化可直接从目标领域证据重塑嵌入几何。我们将这一洞见实例化为DARTopic,这是一个领域无关框架,它从冻结的PLM嵌入构建词级语义图,并联合训练GNN编码器与主题推理模块。在涵盖通用、生物医学和法律领域的三个基准测试中,DARTopic在不进行任何编码器微调的情况下,始终在主题一致性和文档聚类任务上优于强基线,同时展现出对PLM选择的鲁棒性,且相较于基于微调的替代方案具有更优的运行时效率。

英文摘要

Recent advances in neural topic models with pre-trained language models (PLMs) have achieved strong performance by leveraging general-domain pre-training, yet their topic interpretability often degrades on specialized corpora. This limitation primarily stems from the geometry of the embedding space, where domain-specific terms unseen during pre-training collapse into an indistinguishable region, and neither domain-specific re-training, word-level graph enrichment, nor parameter-efficient fine-tuning can restructure this space without inheriting the capacity ceiling of the underlying encoder. Our key insight is that a learnable graph layer operating on token-level PLM embeddings can acquire corpus-specific semantic structure that the frozen encoder lacks, because token-level graphs preserve document-local context that word-level representations discard and joint optimization with the topic objective reshapes embedding geometry directly from target-domain evidence. We instantiate this insight as DARTopic, a domain-agnostic framework that constructs token-level semantic graphs from frozen PLM embeddings and jointly trains a GNN encoder with topic inference. Across three benchmarks spanning general, biomedical, and legal domains, DARTopic consistently outperforms strong baselines in topic coherence and document clus- tering without any encoder fine-tuning, while demonstrating robustness to PLM choice and favorable runtime efficiency over fine-tuning based alternatives.

发表机构

  • TelePIX AI Center(TelePIX人工智能中心)
  • Samsung Electronics(三星电子)
  • School of Computing & Frontier AI Research Centre Macquarie University(麦考瑞大学计算与前沿人工智能研究中心)

机构由 AI 辅助整理,请以论文原文为准。

↑