WPTokens:用于研究维基百科规制的数据集
WPTokens: A Dataset for Studying the Regulation of Wikipedia
浏览论文内容
中文总结 AI 辅助
本文提出WPTokens数据集,基于法语维基百科修订历史,通过两层位置差异追踪内容变更,支持规制研究,涵盖海量数据并配套R包构建贡献者-令牌多重图。
中文摘要 AI 辅助
研究维基百科如何被规制,需要知道谁在何时何地添加、重写、移动和删除了哪些内容。WPTokens是一个数据集,它基于法语维基百科的完整修订历史回答了这些问题。它依赖于两层结构上的位置差异:由句子组成的文本层,以及由wikitext结构元素(模板、链接、参考文献、标题、表格、列表等)组成的节点层。每个内容片段在其整个生命周期内保持一个稳定的标识符,对其进行的每一次更改都被记录为添加、更新、移动、带更新的移动或删除。该数据集基于2026年4月的转储构建,涵盖275万篇文章、1.536亿次修订、3.426亿个令牌和5.737亿次令牌事件,存储在一个索引的SQLite数据库中。一个配套的R包将这些记录转化为时间性的贡献者-令牌多重图,并补充了贡献者之间的回退关系层。我们描述了令牌化、差异算法、数据模型和图构建,并讨论了该方法的局限性。
英文摘要
Studying how Wikipedia is regulated requires knowing who adds, rewrites, moves and deletes which piece of content, when and where. WPTokens is a dataset that answers these questions over the full revision history of the French-language Wikipedia. It relies on a positional diff over two layers: a text layer made of sentences, and a node layer made of the structural elements of wikitext (templates, links, references, headings, tables, lists, etc.). Each piece of content keeps a stable identifier for its whole life, and every change to it is recorded as an addition, an update, a move, a move with update, or a deletion. Built from the April 2026 dump, the dataset covers 2.75 million articles, 153.6 million revisions, 342.6 million tokens and 573.7 million token events, stored in an indexed SQLite database. A companion R package turns these records into temporal contributor--token multigraphs, completed by a layer of revert ties between contributors. We describe the tokenization, the diff algorithm, the data model and the graph construction, and discuss the limitations of the approach.
发表机构
- Universit \'e de Rouen Normandie, DySoLab Mont-Saint-Aignan France
- Universit \'e de Rouen Normandie, DySoLab
机构由 AI 辅助整理,请以论文原文为准。