arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.14491cs.SIcs.CL

人为制造的分裂:剖析七种社交媒体影响行动中的恶意内容

Manufactured Divisiveness: Decomposing the Hostile Content of Seven Social Media Influence Operations

Emilio Ferrara

首次发表
浏览论文内容

中文总结 AI 辅助

研究推特中七种政府认定活动推文里的恶意内容,通过双提示大语言模型检测器和可审计规则区分仇恨与其他分裂形式,发现将所有相关内容计为仇恨会高估,有助于重新定义影响活动及在线话语中仇恨的研究。

中文摘要 AI 辅助

国家支持的影响行动通常被视为“仇恨”和“毒性”的高发来源。我们认为这些比率存在测量误差:背后的检测工具旨在捕捉更广泛的定义,包括针对外群体的敌意或分裂,因此将仇恨过度归因于更应被描述为党派或地缘政治谩骂的内容。我们从推特信息行动档案中七个政府认定的活动的2508万条推文里区分仇恨与其他形式的分裂。先验证了基于双提示大语言模型的检测器,然后制定了可审计规则,将内容分为三个子类别。研究发现将所有这些都报告为仇恨会高估约两倍,只有18.7%是基于身份且有辱人格或煽动性的。七个活动中的六个可分为三种模式,我们称之为“人为制造的分裂”。区分这些结构的界限仍未确定,研究结果有助于在影响活动和更广泛的在线话语背景下重新定义仇恨研究。

英文摘要

State-backed influence operations are routinely measured as high-prevalence sources of ``hate'' and ``toxicity.'' We argue those rates rest on a measurement error: the detectors behind them are validated to catch a broader definition inclusive of hostility or divisiveness aimed at an out-group, and so over-attribute hate to content better described as partisan or geopolitical invective. Across 25.08M tweets from seven government-attributed campaigns in the Twitter Information Operations archive (8,275 accounts), we separate hate from the other forms of divisiveness. We first validate a two-prompt LLM-based detector, matching human labels at Cohen's $κ=0.82$, to identify the broader hostility; we then develop an auditable rule, agreeing with an expert at $κ=0.52$, to further classify this content (5,457 posts) into three sub-categories. About 50.1% are identity-based attacks on people, whereas 30.4% are partisan attacks and 19.5% invective against states and their foreign policy. Reporting all of it as hate therefore overstates hate roughly twofold; only 18.7% is both identity-based and dehumanizing or inciting. Six of seven campaigns sort into three regimes that a single ``hate'' rate flattens, namely identity hate (RU-op and IRA, both Russia-attributed), geopolitical invective (both Iran operations), and partisan divisiveness (both Venezuela operations). We call the shared product $manufactured divisiveness$. The line to separate these constructs itself remains unsettled: on the hardest cases three independent human experts agree only moderately (pairwise $κ=0.37$--$0.50$), and the best of nineteen LLM models tops out at $κ=0.601$ against the experts' majority. Our findings can help redefine the study of hate in the context of influence campaigns and broader online discourse.

发表机构

  • Department of Computer Science University of Southern California(计算机科学系 美国南加州大学)

机构由 AI 辅助整理,请以论文原文为准。

↑