arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.14813cs.CL

超出界限:评估LLM训练数据中极端主义言论的流行程度与内容

Beyond the pale: Assessing prevalence and contents of extremist speech in LLM training data

Dmitry Nikolaev, Ashley A. Mattheis

首次发表
浏览论文内容

中文总结 AI 辅助

本研究针对LLM训练数据的极端主义言论问题,以Dolma语料库为对象,经多步骤提取得出其含数十万份极端主义文档的下限,并探讨相关数据整理与预训练影响。

中文摘要 AI 辅助

尽管研究界对可信且安全的AI话题有浓厚兴趣,但大型语言模型(LLMs)在预训练及后训练阶段所接触的文本语料库构成尚未受到太多关注。本研究探讨LLMs是否会接触到未经过滤、缺乏语境的极端主义言论。研究采用源自官方文件与研究文献的多种极端主义言论定义,结合自动文本处理与专家验证的提取流程,得出支撑OLMo系列模型的开放训练语料库Dolma中极端主义文档流行度的下限。结果显示,Dolma很可能包含数十万份含极端主义内容及多种仇恨言论的文档,其中包括直接呼吁暴力的内容,研究还讨论了这一情况对数据整理及模型预训练的影响。

英文摘要

Despite a strong interest on the part of the research community in the topic of trustworthy and safe AI, the composition of the text corpora that large language models (LLMs) encounter in pre- and post-training has not yet drawn much attention. In this work, we address the question of whether LLMs are exposed to unfiltered, uncontextualised extremist speech. Using several definitions of extremist speech, stemming from official documents and research literature, and an extraction pipeline combining automated text processing with expert verification, we provide a lower bound on the prevalence of extremist documents in Dolma, an open training corpus underpinning the OLMo series of models. We show that Dolma is likely to include hundreds of thousands of documents containing extremist content and hate speech of several types, including direct calls for violence, and discuss the implications of this for data curation and model pre-training.

发表机构

  • University of Manchester(曼彻斯特大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑