从重复到识别:虚假信息叙事模式的归纳发现
From Repetition to Recognition: Inductive Discovery of Disinformation Narratives
- Technische Universität Berlin(柏林工业大学)
- German Research Center for Artificial Intelligence (DFKI)(德国人工智能研究中心)
- BIFOLD – Berlin Institute for the Foundations of Learning and Data(柏林学习与数据基础研究所)
- Johannes Gutenberg-Universität Mainz(美因茨约翰内斯·古腾堡大学)
- Centre for European Research in Trusted AI (CERTAIN)(欧洲可信人工智能研究中心)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对虚假信息叙事挖掘的封闭世界评估局限,提出三层评估框架,比较聚类与图社区方法,发现图方法在开放世界发现中更具优势,并发布人工验证的叙事候选标签。
AI中文摘要:
在虚假信息数据集中,叙事通常被理解为反复出现的解释性模式,这些模式将文本按叙事标签分组。近期工作将叙事挖掘形式化为从语料库中归纳推断叙事标签,但其评估仍依赖于预定义分类体系,这是一种封闭世界设置,无法捕捉参考标签中不存在的叙事。我们引入了一个用于无监督叙事标签生成的三层评估框架:恢复(针对语料库自身的分类体系)、挖掘(针对外部标签集)和发现(无预定义标签)。应用该框架,我们在七个虚假信息数据集上比较了基于聚类和基于图社区的流程,并在其中两个数据集上对发现结果进行了人工验证。在自动化指标下,这两类流程具有互补性,但在一个包含两个突出主题的语料库中,聚类可能将一个主题的生成标签比例降至2%,而基于图的流程则保持平衡。发现验证还揭示了许多单例(源自单一主张的叙事标签,占图输出结果的30-62%),而聚类无法产生这些标签。标注者确认其中许多是可识别的虚假信息叙事,这表明在开放世界发现中,叙事挖掘所假设的重复性可能在语料库之外被识别,而非在语料库之内。我们发布了针对气候阻碍和多元叙事数据集的人工验证叙事候选标签,以支持分类体系开发和数据集扩展。
英文摘要:
In disinformation datasets, narratives are often understood as recurring interpretive patterns that group texts under narrative labels. Recent work formalized narrative mining as inductively inferring narrative labels from corpora, but its evaluation stays tied to predefined taxonomies, a closed-world setting that cannot capture narratives absent from the reference labels. We introduce a three-tier evaluation framework for unsupervised narrative label generation: recovery (against a corpus's own taxonomy), mining (against external label sets), and discovery (without predefined labels). Applying it, we compare clustering-based and graph-community-based pipelines across seven disinformation datasets, with human validation of discovery on two. The two families are complementary under automated metrics, but in a corpus with two prominent topics, clustering can reduce one topic to 2% of generated labels while graph-based pipelines stay balanced. Discovery validation also reveals many singletons (narrative labels derived from single claims, 30-62% of graph outputs), which clustering cannot produce. Annotators confirm many as recognizable disinformation narratives, suggesting that in open-world discovery the repetition assumed by narrative mining may be recognized outside the corpus, not within it. We release human-validated narrative candidate labels for the Climate Obstruction and PolyNarrative datasets to support taxonomy development and dataset extension.