arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.15129cs.CLcs.AI

左分支Transformer在右分支语言中表现优异:数据塑造语言模型的词序偏好

Left-Branching Transformers Excel at Right-Branching Languages: Data Shapes Word Order Preferences in Language Models

Varvara Arzt, Allan Hanbury, Terra Blevins

首次发表
浏览论文内容

中文总结 AI 辅助

研究对比仅解码器Transformer在192种人工及自然语言中的词序偏好,发现其在人工语言呈左分支偏好、自然语言随数据增长偏好SVO,证实词序偏差由数据驱动,或减少语言词序多样性。

中文摘要 AI 辅助

我们系统比较了仅解码器语言模型在192种人工语言和类型多样的自然语言中的词序偏好。在人工语言上,模型表现出左分支偏好,这既不符合自然语言共性,也不符合人类词序学习偏差。在自然语言上,单语模型在小规模时无明确基础词序偏差,但随着数据增长,右分支主-谓-宾(SVO)语言的偏好显现,而尽管SOV是跨语言最常见的词序,其表现却落后。这种SVO优势延伸至多语模型,且与语言资源水平和数据质量相关,而非词序本身。因此,同一架构在人工和自然语言上表现出相反偏好,表明实际观察到的词序偏差是数据驱动的。由于高资源语言绝大多数为SVO,随着大语言模型(LLM)的广泛应用,这些偏差可能逐渐减少词序多样性,尤其是在灵活使用多种词序的语言中。

英文摘要

We systematically compare word order preferences in decoder-only language models across 192 artificial languages and typologically diverse natural languages. On artificial languages, models exhibit a left-branching preference that aligns with neither natural language universals nor human word order learning biases. On natural languages, monolingual models show no clear base word order bias at small scales, but as data grows, a preference for right-branching subject-verb-object (SVO) languages emerges while SOV falls behind despite being the most frequent order cross-linguistically. This SVO advantage extends to multilingual models and correlates with language resource level and data quality rather than word order. Thus, the same architecture exhibits opposite preferences on artificial and natural languages, establishing that word order biases observed in practice are data-driven. Since highly-resourced languages are overwhelmingly SVO, these biases risk gradually reducing word order diversity, particularly in languages that productively use multiple word orders, with the widespread adoption of LLMs.

发表机构

  • Khoury College of Computer Sciences, Northeastern University(东北大学科里计算机科学学院)
  • Faculty of Informatics TU Wien(维也纳技术大学信息学院)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑