arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

合适的工具,合适的工作:仅法语BabyLM的原生语言评估、分词器敏感性与方法论发现

Right Tool, Right Job: Native-Language Evaluation, Tokenizer Sensitivity, and Methodological Findings from a French-Only BabyLM

Adam Zachary Wasserman, David Beauchemin

arXiv 2609.17435首次发表:更新:

发表机构

Open Honest Foundation; Group for Research in Artificial Intelligence of Laval University (GRAIL); Université Laval(开放诚实基金会; 拉瓦尔大学人工智能研究组(GRAIL); 拉瓦尔大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提交仅法语的MéTRON-FR模型至BabyLM,通过原生基准和跨语言评估验证其语法能力,并揭示分词器伪影对零样本评分的影响,提出标准诊断方法。

AI 中文摘要

我们向BabyLM 2026严格赛道提交了MéTRON-FR,一个在9200万法语单词上预训练的1.25亿参数GPT-2模型。它在QFrBLiMP(一个魁北克法语原生语法最小对基准)上得分为85.97%±0.17%,在BabyLM加权排行榜上得分为62.80%。一种结合法语任务数据翻译与秩16 LoRA(低秩适配)的跨语言GLUE(通用语言理解评估)协议产生了明显的任务类型梯度:关系型任务获得可测量的提升,而世界知识型任务则出现退步。双语词汇归纳将法语嵌入与GPT-2对齐,p@1为68.84%±8.61%,是随机水平的18倍,这表明跨语言对齐追踪的是已习得的语法能力而非训练时长。一项消融研究表明,在儿童规模下,单分词零样本评分受分词器和模板伪影主导,这促使将分词器替换敏感性、安慰剂对照提示和原生语言最小对基准作为标准诊断方法。

英文摘要

We submit MéTRON-FR, a 125M GPT-2 pretrained on 92.47M words of French, to the BabyLM 2026 Strict track. It scores 85.97 +/- 0.17% on QFrBLiMP (a native Quebec-French benchmark of grammatical minimal pairs) and 62.80% on the BabyLM-weighted leaderboard. A cross-lingual GLUE (General Language Understanding Evaluation) protocol that combines French task-data translation with rank-16 LoRA (Low-Rank Adaptation) produces a sharp task-type gradient: relational tasks gain measurably, while world-knowledge tasks regress. Bilingual Lexicon Induction aligns the French embeddings to GPT-2 at p@1 = 68.84 +/- 8.61%, 18X above chance, suggesting cross-lingual alignment tracks acquired grammatical competence rather than training duration. An ablation study shows that single-token zero-shot scoring is dominated by tokenizer and template artifacts at the child scale, motivating tokenizer-swap sensitivity, placebo-controlled prompting, and native-language minimal-pair benchmarks as standard diagnostics.

CommentsAccepted at BabyLM Workshop at EMNLP 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑