arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.10525cs.FLcs.LG

极限中的语言生成特征化:有限见证与分离宽度层级

Characterizing Language Generation in the Limit: Finite Witnesses and a Separation-Width Hierarchy

Xiaoyu Li, Andi Han, Jiaojiao Jiang, Junbin Gao

首次发表
浏览论文内容

中文总结 AI 辅助

本文特征化极限中的语言生成,提出有限见证与分离宽度层级,并在Lean中验证,揭示了生成可能性的充分必要条件。

中文摘要 AI 辅助

极限中的语言生成要求从未知无限语言的每一个穷尽正呈现中生成有效的未见元素。我们针对可数宇宙上的任意族来特征化这一任务。当且仅当每个目标可以被赋予一个有限正见证,使得任何有限样本激活的目标具有无限公共交集时,生成才是可能的。必要性方向来自一个通用规范化:通过未确认历史的搜索将任何成功的生成器转换为仅依赖于观察到的集合的生成器。然后我们询问兼容见证必须有多大。正分离宽度记录了最小的统一大小界限,还有两个进一步的级别:无界有限见证和不存在任何兼容的有限见证分配。每个级别都会出现。可数族允许单例见证,显式族实现每个有限宽度,而具有无限公共核心的两个族的并集需要无界有限见证。最后,可数支撑和有限轮廓障碍解释了为什么局部组合数据不能确定极限中的生成。该特征化和完整的宽度层级在Lean中进行了检查,包括简化的规范化和直接的对角捕获引理。随附的Lean开发在此https URL维护。

英文摘要

Language generation in the limit asks for valid unseen elements from every exhaustive positive presentation of an unknown infinite language. We characterize this task for arbitrary families over a countable universe. Generation is possible exactly when each target can be assigned a finite positive witness so that the targets activated by any finite sample have an infinite common intersection. The necessary direction follows from a universal normalization: a search through unconfirmed histories converts any successful generator into one depending only on the observed set. We then ask how large compatible witnesses must be. Positive separation width records the smallest uniform size bound, with two further levels for unbounded finite witnesses and the absence of any compatible finite-witness assignment. Every level occurs. Countable families admit singleton witnesses, explicit families realize every finite width, and a union of two families with infinite common cores requires unbounded finite witnesses. Finally, countable-support and finite-profile obstructions explain why local combinatorial data cannot determine generation in the limit. The characterization and full width hierarchy are checked in Lean, including the simplified normalization and a direct diagonal capture lemma. The accompanying Lean development is maintained at https://github.com/xiaoyulics/language-generation-characterization

发表机构

  • University of New South Wales(新南威尔士大学)
  • University of Sydney(悉尼大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑