首个中文BabyLM挑战:训练数据高效且认知合理的中文语言模型
The First ChineseBabyLM Challenge: training data-efficient and cognitively plausible language models for Chinese
浏览论文内容
中文总结 AI 辅助
首个中文BabyLM挑战将在2026年自然语言处理与中文计算会议举办,要求用1亿中文词元从头训练语言模型,在自然语言理解、认知对齐和汉字知识三轨道评估,不限分词器、模型架构和训练轮数。
中文摘要 AI 辅助
本文介绍了将于2026年自然语言处理与中文计算会议举办的首个中文BabyLM挑战。该挑战要求研究人员用1亿中文词元从头开始训练语言模型,并在自然语言理解、认知对齐和汉字知识三个任务轨道上评估模型。对分词器、模型架构和训练轮数没有限制。挑战详情可在该https网址查看。
英文摘要
This paper presents the first ChineseBabyLM Challenge, organized as part of NLPCC 2026. The challenge asked participants to train language models from scratch using no more than 102M Chinese words. The models were evaluated on three tracks: natural language understanding, cognitive alignment, and Hanzi knowledge. There were no restrictions on tokenizers, model architectures, or the number of training epochs. Eighteen teams submitted 28 distinct models, generating 74 result files. The overall-winning team used a DeBERTa-v2 architecture and introduced an auxiliary pinyin-prediction objective during pretraining. Several submissions also explored curriculum-learning strategies and architectural innovations. Overall, the challenge provides a benchmark for advancing data-efficient and cognitively plausible approaches to Chinese language modeling.
发表机构
- Princeton University(普林斯顿大学)
- Shanghai Jiao Tong University(上海交通大学)
- Chinese Academy of Sciences(中国科学院)
- Columbia University(哥伦比亚大学)
- Beijing Normal University(北京师范大学)
- Tsinghua University(清华大学)
- University of California San Diego(加利福尼亚大学圣地亚哥分校)
- The Hong Kong Polytechnic University(香港理工大学)
机构由 AI 辅助整理,请以论文原文为准。