语法重要吗?一种面向计算社会科学的图增强变分主题模型
Does Syntax Matter? A Graph-Augmented Variational Topic Model for Computational Social Sciences
浏览论文内容
中文总结 AI 辅助
本文提出图增强变分主题模型SCPTM,探究句法在主题建模中的作用,发现句法编码对论证性文本有益,但在信息性语域中引入噪声。
中文摘要 AI 辅助
主题建模在计算社会科学中被广泛用于识别大型文本语料库中的潜在主题。传统方法依赖于词袋表示和诸如LDA之类的生成模型,而最近的方法如BERTopic则基于稠密文档嵌入进行操作。本文提出了结构上下文概率主题模型(SCPTM),这是一种将句法依存关系纳入主题推断的架构。SCPTM将语料库表示为文档和单词的异构图,这些节点通过词汇和句法边连接,并在变分自编码器内通过图注意力网络进行处理,以生成概率性的、混合成员关系的主题分布。我们评估了七种主题建模技术(包括四种SCPTM消融变体),这些技术应用于四个在语域和话语结构上有所不同的语料库。我们的框架结合了连贯性(C_V、C_NPMI)、主题多样性、聚类标签对齐(NMI)以及短语级诊断(互补性和效价差距)。结果表明,SCPTM的神经架构在文档-主题对齐方面相较于生成基线取得了显著提升,但这些提升归因于变分编码器而非语法。语法有助于主题多样性,其中图增强变体在所有语料库中均优于无图基线,并有助于描述符质量:依存路径捕获了审议性语域中的谓词-论元结构和立场,而在技术性和制度性语料库中则被证明是多余的。效价差距在所有变体中均为正,但主要由短语分组而非句法过滤驱动。我们得出结论:句法编码的作用是有条件的——它有益于面向行动、论证性的文本,但在信息性或行政性语域中会引入噪声。
英文摘要
Topic modeling is widely used in computational social sciences to identify latent themes in large text corpora. Traditional approaches rely on Bag-of-Words representations and generative models such as LDA, while recent methods like BERTopic operate on dense document embeddings. This paper introduces the Structural Contextual Probabilistic Topic Model (SCPTM), an architecture that incorporates syntactic dependency relations into topic inference. SCPTM represents a corpus as a heterogeneous graph of documents and words connected by lexical and syntactic edges, processed through a Graph Attention Network within a Variational Autoencoder to produce probabilistic, mixed-membership topic distributions. We evaluate seven topic modeling techniques (including four SCPTM ablations) across four corpora differing in register and discourse structure. Our framework combines coherence (C_V, C_NPMI), topic diversity, clustering-label alignment (NMI), and phrase-level diagnostics (complementarity and valence gap). Results show that SCPTM's neural architecture yields substantial gains in document-topic alignment over generative baselines, but these gains are attributable to the variational encoder rather than to syntax. Syntax contributes to topic diversity, where graph-augmented variants outperform the no-graph baseline across all corpora, and to descriptor quality: dependency paths capture predicate-argument structures and stance in deliberative registers, while proving redundant in technical and institutional corpora. The valence gap is positive across all variants, but driven primarily by phrase grouping rather than syntactic filtering. We conclude that syntactic encoding matters conditionally: it benefits action-oriented, argumentative texts, but introduces noise in informational or administrative registers.
发表机构
- University of Udine(乌迪内大学)
机构由 AI 辅助整理,请以论文原文为准。