FineWeb-CLaR:用于基准对齐语料审计的文化、语言和地区标注
FineWeb-CLaR: Culture, Language, and Region Annotations for Benchmark-Aligned Corpus Auditing
- Saarland University(萨尔兰大学)
- German Research Center for Artificial Intelligence (DFKI)(德国人工智能研究中心)
- Barcelona Supercomputing Center (BSC-CNS)(巴塞罗那超级计算中心)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文提出FineWeb-CLaR,一个基于FineWeb的大规模文化-语言-地区标注数据集,用于审计预训练语料与文化基准的对齐,支持文化现象覆盖范围的直接比较。
AI中文摘要:
语言模型的文化评估覆盖范围和鲁棒性难以诊断,因为预训练语料库和文化基准很少使用可比较的元数据进行索引。基准测试越来越多地针对语言、地区和特定地区实践层面的文化情境现象,而网络规模的语料库通常仅按语言组织。共享的文化-语言-地区层使这些资源具有可比性,从而能够审计目标文化现象是否在预训练数据中有所体现、是否被基准测试评估,或两者兼有。为此,我们引入了FineWeb-CLaR,这是一个从FineWeb和FineWeb-2派生的大规模标注数据集,它将网络文档置于共享的文化-语言-地区轴上,用于语料库审计和基准对齐。FineWeb-CLaR为来自FineWeb和FineWeb-2的完整309亿文档集合标注了基于URL的地区标签和文化主题来源。我们的地区解析器为25.61%的文档(79.2亿)分配了非空地区。对于文化主题分析,我们归纳出特定地区的主题,并将其投影到Liu等人(2025)的文化分类法的14个叶节点上,生成用于语料库侧比较的地区主题分布(LTDs)。我们还使用相同的分类法、语言覆盖范围和地区覆盖范围标注了277个文化NLP基准。这些资源共同实现了语料库侧预训练证据与基准侧评估覆盖范围之间的直接比较。
英文摘要:
Cultural evaluation coverage and robustness in language models are difficult to diagnose because pretraining corpora and cultural benchmarks are rarely indexed with comparable metadata. Benchmarks increasingly target culturally situated phenomena at the level of languages, regions, and locale-specific practices, while web-scale corpora are usually organized only by language. A shared culture-language-region layer makes these resources comparable, enabling audits of whether a target cultural phenomenon is represented in pretraining data, evaluated by benchmarks or both. To this end, we introduce FineWeb-CLaR, a large-scale annotated dataset derived from FineWeb and FineWeb-2 that places web documents on a shared culture-language-region axis for corpus auditing and benchmark alignment. FineWeb-CLaR annotates the full 30.9B-document collection from FineWeb and FineWeb-2 with URL-derived region labels and cultural-topic provenance. Our region resolver assigns a non-empty region to 25.61% of documents (7.92B). For cultural-topic analysis, we induce locale-specific topics and project them onto the 14 leaves of the Cultural Taxonomy of Liu et al. (2025), producing Locale Topic Distributions (LTDs) for corpus-side comparison. We also annotate 277 cultural NLP benchmarks with the same taxonomy, language coverage, and region coverage. Together, these resources enable direct comparison between corpus-side pretraining evidence and benchmark-side evaluation coverage.