发表机构
Stanford University; Cornell University(斯坦福大学; 康奈尔大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文受大语言模型启发,针对语言识别的Gold-Angluin模型相关问题,解决能否用小字母表迹及能否直接从语言定义迹的问题,给出肯定答案,展示了如何定义计算迹实现极限识别,所用字母表与语言定义字母表大小成线性关系且与语言其他属性无关。
AI 中文摘要
受大语言模型的启发,人们对极限情况下语言识别的Gold-Angluin模型重新产生兴趣,希望找到能克服其原始公式负面结果的变体。近期相关论文提出将训练字符串的计算迹和注释作为学习者额外能力的来源。此前工作虽有积极成果,但迹来自显式自动机理论机器模型,其底层令牌词汇量很大。本文解决了两个基本问题:能否用仅小字母表的迹取得积极成果,以及能否直接从语言本身定义迹而无需生成它的底层机器模型。我们对这两个问题都给出了肯定答案:对于任意语言集合,展示了如何定义计算迹以实现极限识别,所用令牌字母表大小与语言定义的字母表大小成线性关系且与语言的其他属性无关。
英文摘要
Motivated by the power of large language models, there has been renewed interest in the Gold-Angluin model of language identification in the limit, with an eye toward variants of the model that might overcome the negative results for its original formulation. Recent papers on this question have proposed looking at computational traces and annotations of training strings as a source of additional power for a learner, reflecting empirical regularities such as the way that commented source code is easier to learn from than arbitrary source code, and text annotated with algorithmically generated chain-of-thought tokens can be easier to learn from than the raw text itself. This recent work has shown positive results for language identification in the presence of such computational traces, but the traces in these positive results come from explicit automata-theoretic machine models that generate the language, where the underlying vocabulary of tokens for the traces is very large. In this paper, we address two fundamental issues left open by this line of work: can we achieve positive results with traces that use only a small alphabet, and can we define traces directly from the language itself, without requiring an underlying machine model that generates it? We establish positive results for both of these questions: for an arbitrary collection of languages, we show how to define computational traces that enable identification in the limit, using an alphabet of tokens that is linear in the size of the alphabet that the languages are defined over, and independent of any other properties of the languages.