AI 中文总结
本文通过人类感知认知证据和多模态模型线索冲突实验,论证语言应位于模型边界与共享码本中,而非内部表征,并提出七点启示。
AI 中文摘要
语言模型基于词元进行计算:语言既是它们的输入,也是它们的输出,并且日益成为它们的内部表征。语言是否应保留所有这些位置,取决于语言对使用它的系统所产生的作用。唯一一个拥有一个世纪相关数据的系统是人类。我们回顾了语言对人类感知、大脑和思维的影响,并针对多模态模型和语言模型审视了相同的证据。在整个过程中,我们将语言视为一种运行在共享码本上的压缩器:一个词是一个索引,内容存在于接收者中,而一个社群维护着该码本。在人类中,这种压缩是可测量的,学习码本会重组感官,而思维在失去语言后仍能存续。随后,我们通过六个视觉语言模型和两个机器人策略上的线索冲突实验,测量了模型在两条线索不一致时所应用的规则。存活的线索按其可靠性规定的权重进行加权,其斜率为理想观察者斜率的11%至82%,且许多答案直接复制了文本。一个策略家族会丢弃一条除其他线索外不增加任何信息的线索,而不是降低其权重;另一个策略家族则保持一个在线索冲突时失效的权重;而一条在每个训练帧中都能识别任务的视觉线索从未被学习到,因为语言通路已经拟合了数据。语言模型是目前人类语言网络的最佳模型,它们已进入人类言语社群,在改变词频的同时,对齐缩小了它们的概念多样性。最后,我们提出了对基于词元的系统的七点启示。语言应位于模型的边界和共享码本中,如同在大脑中一样,而不是作为其内部表征;将码本留在内部的代价是可审计性。
英文摘要
Language models compute over tokens: language is their input, their output, and increasingly their internal representation. Whether language should keep all of these positions depends on what language does to the system that uses it. The one system with a century of data on that question is the human. We review what language does to human perception, the brain, and thought, and read the same evidence against multimodal models and language models. Throughout, we treat language as a compressor that runs on a shared codebook: a word is an index, the content is in the receiver, and a community maintains the codebook. In humans the compression is measurable, learning the codebook reorganizes the senses, and thought survives the loss of language. We then measure the rule that models apply when two cues disagree, with cue-conflict experiments on six vision-language models and two robot policies. Surviving cues are weighted in the order their reliabilities prescribe, at 11 to 82\% of the ideal observer's slope, and many answers copy the text. One policy family drops a cue that adds no information beyond the others rather than down-weighting it, another keeps it at a weight that fails when the cues conflict, and a visual cue that identifies the task in every training frame is never learned, because the language pathway already fits the data. Language models are the best current models of the human language network, and they have entered the human speech community, shifting word frequencies while alignment narrows their conceptual diversity. We close with seven implications for token-based systems. Language belongs at a model's boundary and in the shared codebook, as in the brain, not as its internal representation; the price of leaving the codebook inside is auditability.