arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.11870cs.CL

奥古斯丁婴儿语言模型:直示定义能教给小型语言模型什么,不能教什么

Augustinian BabyLM: What Ostensive Definition Can and Cannot Teach a Small Language Model

  • Institute for Language Sciences(语言科学研究所)
  • Utrecht University(乌得勒支大学)

机构由 AI 辅助整理,请以论文原文为准。

Lisa Bylinina

AI总结:

本文研究直示定义对小型语言模型的影响,通过视觉初始化嵌入,发现其留下持久印记但不影响抽象语法基准,仅在对象属性知识上有优势,并指出现有评估无法捕捉此效应。

AI中文摘要:

语言模型通常以随机词嵌入开始训练:无论'香蕉'意味着什么,都必须从训练语料中学习。我实现了圣奥古斯丁关于词语学习的图景,即通过直示(ostension)来指称意义,用于一个在1000万词上训练的小型掩码语言模型(DeBERTa):在训练之前,视觉上具象的(visually grounded)词元接收由它们所标注的图像区域导出的嵌入;其他词元则随机初始化。视觉初始化留下了可测量的印记,并持续到训练结束。同时,在大多数婴儿语言模型(BabyLM)基准测试中,这种效应仍然不可见,这些基准测试探测的是抽象语法知识:视觉初始化在那里不影响性能。唯一的零样本例外是对象属性知识(COMPS,Misra等人,2023),在那里,种子初始化(seeding)在每个配置中都有帮助。为了跟进这一结果,我构建了一个针对语料定制的视觉属性交换基准(Visual-Property Swap benchmark,Lin等人,2026)版本,该基准测试颜色、材料、大小和形状知识,并包含每个项目的训练频率和种子状态。在这里,视觉种子模型具有持久的、种子复制的优势,且仅限于种子词元。作为因果测试,我展示了先前未种子词元的合成接地(synthetic grounding)将优势精确地转移到这些词元上。功能词和抽象词汇也接收了强大的视觉种子,并在整个训练过程中保留它们,训练目标也利用它们:在每一个种子中,这些词的保留掩码预测损失(held-out mask-prediction loss)都会下降。然而,我运行的任何基准都没有记录到这一点。什么样的评估能捕捉到这一点仍然是一个悬而未决的问题。

英文摘要:

A language model normally begins training with random word embeddings: whatever 'banana' means must be learned from training corpora. I implement St. Augustine's picture of word learning, meaning by ostension, for a small masked language model (DeBERTa) trained on 10M words: before training, visually grounded tokens receive embeddings derived from the image regions they label; other tokens start random. Visual initialization leaves a measurable imprint that lasts until the end of training. At the same time, the effect remains invisible under most BabyLM benchmarks, which probe abstract grammatical knowledge: visual initialization does not affect performance there. The only zero-shot exception is object-property knowledge (COMPS), where seeding helps in every configuration. To follow up on this result, I build a corpus-tailored version of the Visual-Property Swap benchmark, which tests color, material, size, and shape knowledge, with per-item training frequency and seeded status. Here, vision-seeded models have a persistent, seed-replicated advantage. Function words and abstract vocabulary also receive strong visual seeds and retain them throughout training, and the training objective draws on them: held-out mask-prediction loss falls for these words in every seed. However, no benchmark I run registers this. What evaluation would pick this up remains an open question.

↑