arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.10592cs.CLcs.DL

Wieszcz-XIX:一个包含31亿词的1918年前波兰语语料库及从头训练的时间受限语言模型

Wieszcz-XIX: A 3.1-Billion-Word Corpus of Pre-1918 Polish and Temporally Bounded Language Models Trained From Scratch

Szymon Kocur

首次发表
浏览论文内容

中文总结 AI 辅助

本研究构建了31亿词的1918年前波兰语语料库Wieszcz-XIX,从头训练了多规模仅解码器模型,发现其时间受限性优于现代模型,同时发布资源并提示模型存在同期偏见。

中文摘要 AI 辅助

历史波兰语作为一种语言有充分记录,但本文所涵盖时期的机器可读形式标注仅约100万词,其余部分则依赖质量参差不齐的光学字符识别结果。我们提出Wieszcz-XIX,这是一个包含67.5亿个标记(约31亿词)的语料库,涵盖294369份文档,其中大部分是1800年至1918年出版的波兰语期刊,该语料库通过一套过滤、去重、核查1918年后内容泄漏并按文档级别拆分的流程,从Wolne Lektury和互联网档案馆汇编而成。它比同期标注语料库大三个数量级以上,我们对其缺陷进行了量化:识别错误存在假阳性下限,存在近乎相同的重复内容(已被移除),以及1918年后内容泄漏(训练语料库本身已排除,仅在转录源中发现已知残留占其字节数的0.04%至0.38%,因此发布的语料库与训练语料库一一对应)。在人工校正的样本中,文本清晰部分的字符错误率为0.68%,而45%的采样段落无法校正。我们在该语料库上从头训练了一系列仅解码器模型,参数规模从4700万到3.49亿不等,并测量了它们的时间受限性。与两个现代波兰语基础模型(其中一个规模大得多)相比,3.49亿参数模型和与其规模相当的1.07亿参数模型均出现性能交叉;1918年后词汇使它们每字节的成本比同期词汇高约3.1比特,而对比模型未显示此差距;同期词汇使它们的比特成本低于对比模型。在呈现同期文本时,这些模型保留其拼写,而对比模型仅部分保留。增加参数带来的收益约为对数据进行第二次遍历的两倍。我们发布该语料库、代码和权重。内容警告:这些模型会重现同期偏见,包括反犹太主义言论。

英文摘要

Historical Polish is well documented as a language but annotated in machine-readable form only to about a million words for the period this paper covers; the rest sits behind optical character recognition of variable quality. We present Wieszcz-XIX, a corpus of 6.75 billion tokens (about 3.1 billion words) in 294,369 documents, most of them periodical issues, of Polish published from 1800 to 1918, assembled from Wolne Lektury and the Internet Archive by a pipeline that filters, deduplicates, audits for post-1918 leakage and splits at the document level. It is over three orders of magnitude larger than the annotated corpus of the same period, and we quantify its defects: recognition corruption against a false-positive floor, near-identical duplication, which is removed, and post-1918 leakage, which is excluded from the training corpus itself down to a known residue of 0.04 to 0.38% of its bytes, found in the transcribed source, so the published corpus is the trained one document for document. On a hand-corrected sample the character error rate is 0.68% where the text is legible, and 45% of the sampled passages cannot be corrected. On it we train a ladder of decoder-only models from 47M to 349M parameters from scratch, and measure their temporal boundedness. Against two modern Polish base models, one far larger, the 349M shows a crossover, as does the 107M against the comparator of its size: post-1918 vocabulary costs them about 3.1 bits per byte more than period vocabulary, a gap the comparators do not show, and period vocabulary costs them fewer bits than it costs the comparators. Shown period text, the models keep its spelling and the comparators only partly. Adding parameters gains about twice as much as a second pass over the data. We release the corpus, code and weights. Content warning: the models reproduce period prejudice, including antisemitic statements.

补充信息

↑