arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.37408cs.CLcs.SI

看看你让我们聚集了什么:从 Reddit 话语中提取仇恨叙事

Look What You Made Us Cluster: Hate Narrative Extraction from Reddit Discourse

  • DSO National Laboratories(DSO国家实验室)
  • Defence Science and Technology Agency(国防科技局)

机构由 AI 辅助整理,请以论文原文为准。

Annabelle K. L. Chua, Forster J. Khoo, Joel C. R. Tan, Huey Ting Ang, Kheng Hwee Tan, Joel Y. A. Sim, Shirley W. H. Ow, Ria Mundhra, Elsie C. K. Toh, Youfeng Xu, Lynnette H. X. Ng

AI总结:

本文提出一种基于实体-评价对和LLM推理的叙事提取管道,结合Leiden聚类与LLM精炼,从Reddit评论中识别仇恨叙事,并以泰勒·斯威夫特相关数据展示其有效性。

AI中文摘要:

叙事提取使我们能够识别在线的仇恨叙事,支持构建严谨的检测系统。然而,现有的计算方法在精度上受限,因为它们依赖语义表示,往往只能捕捉表面层面的含义。为了检测更精确且可解释的叙事,我们提出了一种提取管道,将叙事表示为实体-评价对。叙事通过一个大型语言模型(LLM)推理过程提取,该过程扩展了基于方面的情感分析,识别方面,将其判断类型分类为评价的基础,并据此推导出评价。提取的叙事随后使用 Leiden 算法进行聚类,之后通过 LLM 引导的精炼过程将聚类解析到预期的粒度水平。我们用 2024 年批评泰勒·斯威夫特的英文 Reddit 评论来展示这一叙事管道,分析一个表现出仇恨言论模式的代表性聚类,以证明其解释价值。

英文摘要:

Narrative extraction allows us to identify online hate narratives, supporting the construction of rigorous detection systems. Existing computational approaches, however, are limited in precision as they rely on semantic representations, which tend to capture only surface-level meaning. To detect more precise and interpretable narratives, we present an extraction pipeline that represents narratives as entity-evaluation pairs. Narratives are extracted using a Large Language Model (LLM) reasoning process that extends Aspect-Based Sentiment Analysis, identifying the aspect, classifying its judgement type as the basis for evaluation, and deriving the evaluation accordingly. Extracted narratives are then clustered using Leiden, following which clusters are resolved to an intended level of granularity through an LLM-guided refinement process. We illustrate this narrative pipeline with English Reddit comments from 2024 that criticize Taylor Swift, analyzing a representative cluster that exhibits hate speech patterns to demonstrate its interpretive value.

补充信息

↑