从文本语料库学习人类语言的形式化局限
A Formal Limitation on Learning Human Language From Textual Corpora
浏览论文内容
中文总结 AI 辅助
该研究从信息论角度推导了听者从话语表示恢复说话者意图含义的概率上界,通过三类实验验证了理论,表明任何文本特征提取器的表示都无法超越该上界。
中文摘要 AI 辅助
听者能否仅通过话语的形式恢复说话者的意图含义?我们从信息论角度回答该问题,针对任意文本特征提取器(包括当代大语言模型的隐藏状态)给定的听者展开研究。将语言使用建模为含义、语境与话语的联合分布,我们推导了解码器从话语表示中恢复说话者意图含义的概率上界。这些上界由形式对含义的不确定性决定,该不确定性可分为不可约部分与仅能由(非语言的)语境、而非仅话语本身解决的部分。由于这些量是语言的固有属性,无论使用多少文本或监督数据生成何种表示,都无法超越这些上界;无论含义空间是离散还是连续,这些上界均成立。在人工语言、汉语零代词消解和颜色指称任务上的实验为该理论提供了实证支持。
英文摘要
Can a listener recover what a speaker means from the form of an utterance alone? We answer this question information-theoretically, and for a listener given by any featurizer of text, including the hidden states of contemporary large language models. Modeling language use as a joint distribution over meanings, contexts, and utterances, we derive upper bounds on the probability that a decoder recovers a speaker's intended meaning from a representation of the utterance. The bounds are governed by the uncertainty that form leaves about meaning, which splits into an irreducible part and a part that only (extralinguistic) context, but never the utterance alone, can resolve. Because these quantities are intrinsic to language, no representation, however much text or supervision produced it, can surpass them. The bounds apply, moreover, to meaning spaces that are discrete or continuous. We provide empirical evidence in support of the theory through experiments on artificial languages, Mandarin zero-pronoun resolution, and color reference.
发表机构
- Universitat Pompeu Fabra(庞培法布拉大学)
- ETH Zürich(苏黎世联邦理工学院)
机构由 AI 辅助整理,请以论文原文为准。