共同语言学:人工智能增强的语言学理论构建
Co-Linguistics: AI-augmented Theory Construction in Linguistics
浏览论文内容
中文总结 AI 辅助
本文提出将AI作为共同科学家用于语言学理论构建与评估,通过使理论明确化、比较竞争理论及提出新理论来加速研究,同时强调人类在提供方向和评估中的核心作用。
中文摘要 AI 辅助
近年来,大型语言模型(LLMs)在语言学中被研究,作为人类语言能力的潜在模型。在此,我们讨论人工智能的一种完全不同的用途,即作为共同科学家,帮助构建和评估语言学理论(我们将这一结果称为“共同语言学”)。自20世纪60年代以来,语言学发展的理论原则上可以在数学上形式化,通常采用形式语言理论或模型理论的语言。因此,数学领域的人工智能革命将对语言学产生影响——但有一个关键的转折:证明新定理很少是语言学家的目标。相反,人们寻求找到最佳的公理集来推导经验陈述。人工智能可以通过使现有理论完全明确化、比较竞争理论,以及更雄心勃勃地提出新理论(在机器学习中,这与“程序归纳”相关)来加速研究。它还将通过加速关键预测的识别和测试来帮助评估理论,这得益于对数据的无与伦比的访问(在机器学习中,这与“主动学习”相关)。虽然从理论评估到理论构建的循环可能引发语言学理论的递归且可能自主的改进,但人类仍然处于核心地位:语言学家提供科学方向并从概念上评估理论,实验参与者需要评估超出LLMs范围的经验预测。
英文摘要
LLMs have been studied in recent linguistics as potential models of humans' linguistic abilities. Here we discuss an entirely different use of AI, namely as a co-scientist, to help construct and assess linguistic theories (we refer to the result as "Co-Linguistics"). Since the 1960s, linguistics has developed theories that are in principle mathematically formalizable, often in the language of formal language theory or model theory. The AI revolution in mathematics will thus have consequences in linguistics-but with an essential twist: proving new theorems is rarely the linguist's goal. Rather, one seeks to find the best set of axioms to derive empirical statements. AI could accelerate research by making existing theories fully explicit, by comparing competing theories, and more ambitiously, by proposing new theories (in machine learning, this relates to "program induction"). It will also help assess theories by accelerating the identification and test of crucial predictions, thanks to unparalleled access to data (in machine learning, this relates to "active learning"). While the cycle from theory evaluation to theory construction may give rise to recursive and possibly autonomous improvement of linguistic theories, humans remain central: linguists provide scientific directions and evaluate theories conceptually, and experimental participants are needed to assess empirical predictions that are outside the reach of LLMs.