实体追踪在亚十亿参数规模的语言模型中出现,且在自然叙事中超越人类表现
Entity tracking emerges in sub-billion parameter language models and exceeds human performance in naturalistic narratives
浏览论文内容
中文总结 AI 辅助
本研究发现4.1亿参数的语言模型已具备人类水平的实体追踪能力,且随规模提升而改善,远超人类表现,该能力在远小于预期的模型规模时出现。
中文摘要 AI 辅助
理解语言需要追踪语篇中的实体,即明确事物的位置及其变化,即便这些信息未被明确表述。语言模型是否以类似人类的方式执行此类追踪尚不明确,部分原因在于现有评估依赖的是人工任务,与自然语言理解相去甚远,且缺乏与人类的对比。本研究使用不同复杂程度的自然叙事,对语言模型和人类(N=48)的实体追踪能力进行评估。在人类中,研究发现实体追踪能力会随叙事复杂度的提升而下降,但与叙事长度无关;在语言模型中,研究发现4.1亿参数规模的模型已具备人类水平的实体追踪能力,远低于先前研究确定的数十亿参数的代码专用模型,且该能力会随模型规模提升而改善,当前模型的表现远超人类。综合来看,这些结果表明,作为语言理解核心组成部分的实体追踪能力,在远小于此前认为的模型规模时就已出现。
英文摘要
Understanding language requires tracking entities across discourse - i.e., knowing where things are and how they change, even when not explicitly stated. Whether language models perform such tracking in a human-like fashion remains unclear, in part because existing evaluations rely on artificial tasks, far removed from natural language comprehension, and lack comparisons to humans. Here, we evaluate entity tracking in both language models and humans (N = 48) using naturalistic narratives at multiple levels of complexity. In humans, we find that entity tracking degrades specifically with narrative complexity, not narrative length. In language models, we find that human-level entity tracking is already present at 410 million parameters - well below the multi-billion parameter, code-specialised models identified by prior work - and improves with scale, with contemporary models far exceeding human performance. Together, these results demonstrate that entity tracking, a core component of language understanding, emerges at model scales far smaller than previously thought.
发表机构
- IDEAS Research Institute(IDEAS研究所)
- Max Planck Institute for Psycholinguistics(马克斯·普朗克心理语言学研究所)
- University of Amsterdam(阿姆斯特丹大学)
机构由 AI 辅助整理,请以论文原文为准。