arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

统计文档数量不等于统计文本数量:网络PDF语料库统计中的单位偏差

Counting Documents Is Not Counting Text: Unit Bias in Web-PDF Corpus Statistics

Luca Foppiano

arXiv 2608.16390首次发表:更新:

发表机构

Common Crawl Foundation(Common Crawl基金会)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究指出网络PDF语料库统计存在单位偏差,其以文档为单位的统计与标记数差异显著,Common Crawl的截断上限已造成大量文本丢失,建议同时用文档和标记单位报告语料库统计数据。

AI 中文摘要

PDF语料库以标记数(tokens)来宣传自身规模,但它们发布的所有比率(覆盖范围、OCR路由、重新获取恢复率、语言构成)都是按文档计算的,且没有任何语料库分解其总标记数。这两种统计单位的差异极为显著。在CC-MAIN-2021-31-PDF-UNTRUNCATED(包含790万份网络PDF,共326亿个标记)中,3.02%的含文本文档拥有一半的标记数(基尼系数为0.807);超过50页的文档占语料库的5.00%,却承载了53.53%的文本;由TeX工具链生成的PDF占文档总数的1.66%,却占文本总量的4.05%。最明显的受害者是Common Crawl的截断上限:它影响了23.06%的文档和63.08%的文本。重建截断文件并提取两个版本后,两个广泛使用的库分别恢复了11.4%和1.4%的该部分文本;72%至97%的受影响文档无法提取到任何内容;语料库约55%至62%的文本因此丢失。在2025年3月采用的5 MiB上限下,仍会有30.19%的标记被截断,且这些文档的恢复率仅从3.3%升至13.2%。作者建议语料库统计应同时以文档和标记两种单位报告。

英文摘要

PDF corpora advertise their size in tokens, but every rate they publish (coverage, OCR routing, re-fetch recovery, language mix) is computed per document, and none decomposes its token total. Because PDF length is extremely skewed, the two units can describe the same corpus very differently. We ask how the headline statistics of a web-PDF corpus change when each document is weighted by the text it contributes rather than counted once. We used CC-MAIN-2021-31-PDF-UNTRUNCATED (7.9M Common Crawl PDFs, 32.6B tokens), the one public corpus that pairs the fragments Common Crawl stored with the re-fetched originals. Text mass is highly concentrated: 3.02% of text-bearing documents hold half the tokens (Gini 0.807). The clearest consequence is Common Crawl's payload cap, which truncated 23.06% of these documents but 63.08% of their text. Reconstructing the truncated fragments and extracting both versions, two widely used text-layer parsers recover only 1.4% and 11.4% of that exposed text, so roughly 55-62% of the corpus's text is unrecoverable from the crawl by such pipelines; under the 5MiB cap adopted in March 2025, 30.19% of tokens would still be exposed. We recommend that corpus statistics be reported in both units, documents and tokens.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑