AI 中文总结
TatBLiMP 是首个鞑靼语语言学最小对基准,覆盖16种形态句法现象,通过单语素扰动构建句子对,评估语言模型语法性,发现针对性训练比参数规模更关键。
AI 中文摘要
我们推出了 TatBLiMP,这是首个针对鞑靼语(tt,ISO 639-3 tat)的语言学最小对基准,鞑靼语是一种使用西里尔字母书写的钦察突厥语。据我们所知,这是对鞑靼语语言模型进行的首次语法性评估,因为即使是包含 101 种语言的 MultiBLiMP 也不包含鞑靼语。TatBLiMP 涵盖 16 种形态句法现象,共 1248 个句子对。每对句子仅在一个语素上有所不同,一个语法正确,一个语法错误。当模型为语法正确的句子分配更高概率时,即视为通过该句子对。评分比较模型已分配的概率,因此该基准无需文本生成,也无需解析器,可在基础模型和中期训练检查点上运行。TatBLiMP 将 TurBLiMP 的现象清单和单语素拆分操作适配到鞑靼语,并增加了一种鞑靼语特有的现象:数词和量词后的裸名词数。每个句子对中语法正确的句子均来自鞑靼语文学散文的实证句子。语法错误的句子由 apertium-tat 转换器通过确定性单语素扰动生成。每个句子对均经母语者认可。构建遵循合理性原则,因此语法错误的句子是合理的现实错误,而非任意破坏。在从头训练的鞑靼语模型、跨语言适配模型和前沿多语言大语言模型中,该基准追踪的是针对性的鞑靼语训练,而非参数规模。一个 4.78 亿参数的从头训练模型和一个 1.25 亿参数的单语模型领先,准确率接近 0.97;一个 70 亿参数的适配模型落后;参数规模在 300 亿至 1200 亿之间的前沿大语言模型准确率降至 0.80 至 0.92;一个轻量微调的多语言模型表现最弱。最后,我们指出该基准的主要局限性:其继承的分类体系遗漏了母语者最关注的形态音位学、元音和谐和辅音同化,我们勾勒了一个可补充这些内容的原生第二层。
英文摘要
We introduce TatBLiMP, the first benchmark of linguistic minimal pairs for Tatar (tt, ISO 639-3 tat), a Qypchaq Turkic language written in Cyrillic. To our knowledge it is the first grammaticality evaluation for Tatar language models of any kind, since even the 101-language MultiBLiMP does not include Tatar. TatBLiMP covers 16 morphosyntactic phenomena in 1248 sentence pairs. Each pair differs by a single morpheme, one grammatical and one ungrammatical. A model passes a pair when it assigns higher probability to the grammatical member. Scoring compares probabilities the model already assigns, so the benchmark needs no text generation and no parser, and it runs on base models and on mid-training checkpoints. TatBLiMP adapts the phenomenon inventory and single-morpheme breaking operations of TurBLiMP to Tatar and adds one phenomenon specific to Tatar, bare-noun number after numerals and quantifiers. The grammatical member of every pair is an attested sentence from Tatar literary prose. The ungrammatical member is produced by a deterministic single-morpheme perturbation with the apertium-tat transducer. Every pair is ratified by a native speaker. A plausibility principle governs construction, so the ungrammatical member is a plausible real-world error rather than an arbitrary corruption. Across from-scratch Tatar models, cross-lingual adaptations, and frontier multilingual LLMs, the benchmark tracks focused Tatar training rather than parameter scale. A 478M from-scratch model and a 125M monolingual model lead near 0.97, a 7B adaptation trails, frontier LLMs of 30-120B parameters fall to 0.80-0.92, and a lightly tuned multilingual model is weakest. We close with the benchmark's main limitation. Its inherited taxonomy omits the morphophonology, vowel harmony and consonant assimilation, that is most salient to native speakers, and we sketch a native second layer that would add it.
Comments11 pages. Dataset: https://huggingface.co/datasets/ilchats/TatBLiMP