SWARM:一个用于搜索引擎结果中俄语宣传检测的多语言人工标注数据集
SWARM: A Multilingual Human-Annotated Dataset for Russian Propaganda Detection in Search Engine Results
浏览论文内容
中文总结 AI 辅助
本文提出多语言数据集SWARM,含2,183条搜索引擎结果,用于检测俄语宣传,发现内容级分析优于来源屏蔽列表,最强LLM的F1达0.73。
中文摘要 AI 辅助
俄罗斯国家宣传跨越多种语言和在线空间传播。然而,大多数计算研究仅考察其中一个空间,通常是社交媒体,且仅涉及一两种语言,并分析来源而非内容。我们引入了SWARM(多语言俄语宣传标注的搜索网页文档),这是一个包含2,183条搜索引擎结果的数据集,涵盖九种语言和多样化的网页领域(如新闻、博客、政府网站),每条结果均由训练有素的编码员标注其是否支持反复出现的俄罗斯宣传叙事。我们针对这些标签基准测试了基于来源的屏蔽列表、监督分类器和零样本大语言模型。屏蔽列表漏掉了大多数支持宣传的文档,因为此类内容不仅限于被标记的“宣传”媒体,也出现在主流媒体上。内容级分析有所帮助,但效果取决于模型:最强的大语言模型达到了正类F1分数0.73,而监督分类器仅达到约0.5,较小的语言模型过度预测支持,将主题相关性误认为背书。因此,检测搜索传播的宣传需要针对每种语言进行内容级评估,我们希望SWARM及我们的评估代码能够实现这一点。
英文摘要
Russian state propaganda spreads across many languages and online spaces. Yet, most computational work examines only one such space, usually social media, in one or two languages, and analyses sources rather than content. We introduce SWARM (Search-Web documents Annotated for Russian propaganda, Multilingual), a dataset of 2,183 search engine results across nine languages and diverse web domains (e.g., news, blogs, government sites), each annotated by trained coders for whether it supports a recurring Russian propaganda narrative. We benchmark a source-based blocklist, supervised classifiers, and zero-shot LLMs against these labels. The blocklist misses most propaganda-supporting documents, because such content is not confined to flagged "propaganda" outlets but also appears on mainstream ones. Content-level analysis helps, though how much depends on the model: the strongest LLM reaches a positive-class F1 of 0.73, whereas the supervised classifiers reach only about 0.5, with the smaller LLMs over-predicting support, mistaking topical relevance for endorsement. Detecting search-borne propaganda thus requires per-language, content-level evaluation, which we hope SWARM and our evaluation code enable.
发表机构
- Weizenbaum Institute(魏茨鲍姆研究所)
- University of Oxford(牛津大学)
- University of Bern(伯尔尼大学)
机构由 AI 辅助整理,请以论文原文为准。