发表机构
NOVA School of Science and Technology; NOVA LINCS; Fundação para a Ciência e Tecnologia; Instituto Superior Técnico, Universidade de Lisboa; Instituto de Telecomunicações(NOVA科学与技术学院; NOVA LINCS; 科学技术基金会; 里斯本大学高等技术学院; 电信研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出高效流水线,从411TB葡萄牙语网络数据中构建410亿词元的高质量欧洲葡萄牙语语料库,通过抓取后处理等创新使产量提升19.04%,优化LLM预训练。
AI 中文摘要
为欧洲葡萄牙语(PT-PT)等区域语言变体构建网络语料库,主要受方言重叠(尤其是与巴西葡萄牙语PT-BR的重叠)和数据处理的规模所限制。本文提出了一种高效的流水线,用于从葡萄牙语网络中整理出可用于生产的PT-PT语料库,该语料库涵盖来自此http URL的411 TB原始数据。我们引入了一种新颖的抓取后处理模块,在过滤之前移除样板文本和重复行。这种早期干预通过挽救标准启发式过滤器过早丢弃的有效文本,使最终文档产量提高了19.04%。结合严格的语言识别、加权模糊去重和神经质量分类,我们的流水线提供了一个可扩展的框架,以及一个针对大语言模型预训练优化的干净、具有代表性的语料库。
英文摘要
Curating Web corpora for regional language variants like European Portuguese (PT-PT) is heavily bottlenecked by dialectal overlap (mainly with PT-BR) and data processing scale. This paper presents an efficient pipeline to curate a production-ready PT-PT corpus from the Portuguese Web, spanning 411 TB of raw data from Arquivo.pt. We introduce a novel post-scraping block that removes boilerplate and line duplicates prior to filtering. This early-stage intervention increases final document yield by 19.04% by rescuing valid text that standard heuristic filters prematurely discard. Integrated with rigorous language identification, weighted fuzzy deduplication, and neural quality classification, our pipeline offers a scalable framework and a clean, representative corpus optimized for LLM pre-training.
Comments16 pages, 9 figures, EMNLP 2026 Main