大规模归纳式声明提取
Inductive Claims Extraction at Scale
浏览论文内容
中文总结 AI 辅助
本文提出一种利用大型语言模型从社交媒体大规模语料中归纳式提取并编目声明的流程,并在两个Twitter数据集上验证其有效性,为计算社会科学研究提供新工具。
中文摘要 AI 辅助
社交媒体上很大一部分政治话语是在声明层面上构建和表达的:即陈述性的、通常是单句的表述,它们传达了对现实的特定解读,范围可以从事实性到评价性。此外,声明并非随机出现,而是会聚合、以模式重复出现,并与不同的世界观相关联。当与社会网络分析等结构性计算工具配对时,声明可以成为研究政治现象(如回音室或两极分化)的有力分析单元。在本文中,我们提出了一个使用大型语言模型(LLM)从大规模社交媒体语料库中归纳式提取和编目声明的流程,并将其应用于两个不同的Twitter数据集:一个与2020年美国总统选举相关,另一个与2022年FIFA世界杯相关。我们通过将流程的召回率和精确率与人工标注样本进行比较来全面评估该方法,进行消融研究以隔离其各组成部分的贡献,并进行定性错误分析。我们在计算社会科学研究的背景下讨论了该方法的价值,并通过展示从每个数据集获得的声明目录来说明其能力。
英文摘要
A large part of political discourse on social media is built and expressed at a level of claims: i.e. declarative, typically single-clause statements, which convey a particular interpretation of reality and can range from factual to evaluative. Moreover, rather than occurring randomly, claims coalesce, recur in patterns, and come to be associated with different world views. When paired with structural computational tools such as Social Network Analysis, claims can be a powerful unit of analysis to study political phenomena such as echo chambers or polarisation. In this paper, we present a pipeline that uses a large language model (LLM) to inductively extract and catalogue claims from large social media corpora, and apply it to two different Twitter datasets: one relating to the 2020 US presidential election and the other to the 2022 FIFA World Cup. We comprehensively evaluate the approach by measuring the pipeline's recall and precision against manually annotated samples, run ablation studies isolating the contribution of its various components, and perform a qualitative error analysis. We discuss the value of the approach in the context of Computational Social Science research, and illustrate its capabilities by presenting the claims catalogue obtained from each dataset.
发表机构
- University of Edinburgh(爱丁堡大学)
机构由 AI 辅助整理,请以论文原文为准。