arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Fine PT-PT Web:高质量410亿词元的欧洲葡萄牙语网络数据集

Fine PT-PT Web: A High-Quality 41 Billion Tokens Data Collection of the European Portuguese Web

Gonçalo Vinagre, Rui Pedro Guerra, Pedro Gomes, Miguel Moura Ramos, Duarte Miguel Alves, Afonso Simplício, Diogo Tavares, David Semedo, Daniel Gomes, João Magalhães

arXiv 2609.07699首次发表:更新:

发表机构

NOVA School of Science and Technology; NOVA LINCS; Fundação para a Ciência e Tecnologia; Instituto Superior Técnico, Universidade de Lisboa; Instituto de Telecomunicações(NOVA科学与技术学院; NOVA LINCS; 科学技术基金会; 里斯本大学高等技术学院; 电信研究院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出高效流水线,从411TB葡萄牙语网络数据中构建410亿词元的高质量欧洲葡萄牙语语料库,通过抓取后处理等创新使产量提升19.04%,优化LLM预训练。

AI 中文摘要

为欧洲葡萄牙语(PT-PT)等区域语言变体构建网络语料库,主要受方言重叠(尤其是与巴西葡萄牙语PT-BR的重叠)和数据处理的规模所限制。本文提出了一种高效的流水线,用于从葡萄牙语网络中整理出可用于生产的PT-PT语料库,该语料库涵盖来自此http URL的411 TB原始数据。我们引入了一种新颖的抓取后处理模块,在过滤之前移除样板文本和重复行。这种早期干预通过挽救标准启发式过滤器过早丢弃的有效文本,使最终文档产量提高了19.04%。结合严格的语言识别、加权模糊去重和神经质量分类,我们的流水线提供了一个可扩展的框架,以及一个针对大语言模型预训练优化的干净、具有代表性的语料库。

英文摘要

Curating Web corpora for regional language variants like European Portuguese (PT-PT) is heavily bottlenecked by dialectal overlap (mainly with PT-BR) and data processing scale. This paper presents an efficient pipeline to curate a production-ready PT-PT corpus from the Portuguese Web, spanning 411 TB of raw data from Arquivo.pt. We introduce a novel post-scraping block that removes boilerplate and line duplicates prior to filtering. This early-stage intervention increases final document yield by 19.04% by rescuing valid text that standard heuristic filters prematurely discard. Integrated with rigorous language identification, weighted fuzzy deduplication, and neural quality classification, our pipeline offers a scalable framework and a clean, representative corpus optimized for LLM pre-training.

Comments16 pages, 9 figures, EMNLP 2026 Main

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑