发表机构
School of Computing, Institute of Science Tokyo(科学技术学院计算学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
FreqBLiMP通过频率控制的最小对基准,揭示大语言模型在词汇稀有性下语法判断总体稳健但特定现象上脆弱。
AI 中文摘要
最小对基准(如BLiMP)通过测试语言模型(LMs)是否偏好可接受的句子而非最小差异的不可接受句子来评估语言知识。然而,这些基准在很大程度上忽略了词汇频率变化,尽管词汇频率是自然语言使用中普遍且高度偏斜的属性。因此,现有评估并未测试当对比涉及稀有词汇项时,语法偏好是否保持稳定。我们提出了FreqBLiMP,这是BLiMP的一个频率控制扩展,它在明确的齐普夫频率机制下重新生成所有67个范式,同时保留每个最小对的语法对比。我们评估了多个跨规模的开权重LLM家族,发现词汇频率的降低会产生句子似然性的一致、单调下降,但整体对比可接受性准确率仅适度降低。然而,这种总体稳定性掩盖了不同语言现象间的显著差异,LLMs在显性形态句法泛化上保持稳健,而在需要词元特定信息的现象上则性能下降。
英文摘要
Minimal-pair benchmarks such as BLiMP evaluate linguistic knowledge by testing whether language models (LMs) prefer acceptable sentences over minimally different unacceptable ones. However, these benchmarks largely ignore lexical frequency variation, despite lexical frequency being a pervasive and highly skewed property of natural language use. Consequently, existing evaluations do not test whether grammatical preferences remain stable when contrasts involve rare lexical items. We introduce FreqBLiMP, a frequency-controlled extension of BLiMP that regenerates all 67 paradigms under explicit Zipf-frequency regimes while preserving each minimal-pair's grammatical contrast. Evaluating multiple open-weight LLM families across scales, we find that decreasing lexical frequency produces a consistent, monotonic decrease in sentence likelihood, but only a modest reduction in overall contrastive acceptability accuracy. However, this aggregate stability masks substantial variation across linguistic phenomena, with LLMs remaining robust on overt morphosyntactic generalization while degrading on phenomena that require lemma-specific information.
CommentsAccepted to EMNLP 2026 Main Conference. Repo link: https://github.com/TimeTravelerTy/freqblimp-generation