arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.11606cs.CLcs.DScs.LG

用于语言识别的全局一致着色方案

Globally Consistent Coloring Schemes for Language Identification

Moses Charikar, Jon Kleinberg, Chirag Pabbaraju

首次发表
浏览论文内容

中文总结 AI 辅助

研究对抗性语言学习所需额外信息,探讨终端着色能否替代整个颜色序列用于语言识别,证明每个字符串只需一个终端位,全局构造用超限递归,还表明有限颜色的博雷尔映射终端着色不能识别所有可数子集合。

中文摘要 AI 辅助

我们研究了进行对抗性语言学习所需的额外信息有多少。在戈德的极限语言识别模型中,学习者会得到来自可数语言集合中一种未知语言的字符串枚举。学习者在枚举过程中猜测语言的身份,若最终所有猜测都是正确语言则成功。经典结果表明许多自然集合无法以此方式学习。近期关于追踪着色的工作克服了这一障碍。我们探讨学习者是否真的需要整个颜色序列,还是每个字符串末尾的一种颜色(终端着色)就足以进行语言识别。我们证明每个可数的无限语言集合,每个字符串只需一个终端位就足够。实际上,着色可以独立于集合选择:有一种对每个无限语言的双色终端着色的单一分配方式,使得相同的预先分配的着色能识别每个可数子集合。我们的全局构造使用超限递归,并且证明对于任何有限数量的颜色,这种非构造性是不可避免的。作为一种构造性概念,我们使用博雷尔映射的形式主义;我们表明由博雷尔映射定义的具有有限数量颜色的全局终端着色不能识别所有可数子集合。相比之下,已知的追踪着色构造编码为终端着色时是博雷尔的,但需要无限多种颜色。

英文摘要

We study how little extra information is needed to make adversarial language learning possible. In Gold's model of language identification in the limit, a learner is given an enumeration of the strings from an unknown language chosen from a countable language collection. The learner guesses the identity of the language over the course of the enumeration, and it succeeds if, eventually, all of its guesses are the correct language. Classical results of Gold and Angluin show that many natural collections cannot be learned in this way. Recent work on trace colorings, motivated by the success of thinking-trace strategies in language learning, overcomes this obstruction by annotating every symbol of every string with a color. We ask whether the learner really needs this whole sequence of colors, or whether one color at the end of each string (a terminal coloring) is enough for language identification. We show that just one terminal bit per string is enough for every countable collection of infinite languages. In fact, the colorings can be chosen collection-independently: there is a single assignment of a two-color terminal coloring to every infinite language such that the same preassigned colorings identify every countable subcollection. Thus, in this model, an entire color trace can be compressed to one bit attached to the end of each example. Our global construction uses transfinite recursion, and we prove that this kind of nonconstructivity is unavoidable for any bounded number of colors. As a notion of constructivity, we use the formalism of Borel maps (a regularity condition satisfied by natural explicit constructions); we show that no global terminal coloring with a finite number of colors defined by a Borel map can identify all countable subcollections. By contrast, known trace-coloring constructions are Borel when encoded as terminal colorings, but require infinitely many colors.

发表机构

  • Stanford University(斯坦福大学)
  • Cornell University(康奈尔大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑