发表机构
Open Honest Foundation; Group for Research in Artificial Intelligence of Laval University (GRAIL); Université Laval(开放诚实基金会; 拉瓦尔大学人工智能研究组(GRAIL); 拉瓦尔大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提交仅法语的MéTRON-FR模型至BabyLM,通过原生基准和跨语言评估验证其语法能力,并揭示分词器伪影对零样本评分的影响,提出标准诊断方法。
AI 中文摘要
我们向BabyLM 2026严格赛道提交了MéTRON-FR,一个在9200万法语单词上预训练的1.25亿参数GPT-2模型。它在QFrBLiMP(一个魁北克法语原生语法最小对基准)上得分为85.97%±0.17%,在BabyLM加权排行榜上得分为62.80%。一种结合法语任务数据翻译与秩16 LoRA(低秩适配)的跨语言GLUE(通用语言理解评估)协议产生了明显的任务类型梯度:关系型任务获得可测量的提升,而世界知识型任务则出现退步。双语词汇归纳将法语嵌入与GPT-2对齐,p@1为68.84%±8.61%,是随机水平的18倍,这表明跨语言对齐追踪的是已习得的语法能力而非训练时长。一项消融研究表明,在儿童规模下,单分词零样本评分受分词器和模板伪影主导,这促使将分词器替换敏感性、安慰剂对照提示和原生语言最小对基准作为标准诊断方法。
英文摘要
We submit MéTRON-FR, a 125M GPT-2 pretrained on 92.47M words of French, to the BabyLM 2026 Strict track. It scores 85.97 +/- 0.17% on QFrBLiMP (a native Quebec-French benchmark of grammatical minimal pairs) and 62.80% on the BabyLM-weighted leaderboard. A cross-lingual GLUE (General Language Understanding Evaluation) protocol that combines French task-data translation with rank-16 LoRA (Low-Rank Adaptation) produces a sharp task-type gradient: relational tasks gain measurably, while world-knowledge tasks regress. Bilingual Lexicon Induction aligns the French embeddings to GPT-2 at p@1 = 68.84 +/- 8.61%, 18X above chance, suggesting cross-lingual alignment tracks acquired grammatical competence rather than training duration. An ablation study shows that single-token zero-shot scoring is dominated by tokenizer and template artifacts at the child scale, motivating tokenizer-swap sensitivity, placebo-controlled prompting, and native-language minimal-pair benchmarks as standard diagnostics.
CommentsAccepted at BabyLM Workshop at EMNLP 2026