arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.18305cs.LGcs.AI

信息阴影:衡量语言模型学习的结构限制

The Information Shadow: Measuring Structural Limits on What Language Models Can Learn

发表机构Sirena人工智能公司
查看机构详情
  • Sirena Ai(Sirena人工智能公司)

机构由 AI 辅助整理,请以论文原文为准。

Priyansh Srivastava, Romit Chatterjee

首次发表
浏览论文内容

中文总结 AI 辅助

研究语言模型学习的结构限制,引入信息阴影概念,包含语言无法表达的结构等三类。通过语言压缩残差等探针测试,揭示不同类型限制,发布探针套件并探讨其对基准设计等方面的影响。

中文摘要 AI 辅助

语言模型知识的某些限制并非数据覆盖的差距,而是从文本学习的结构属性。我们引入了信息阴影,即文本训练的学习者无论规模大小都无法获取的现象区域,包括:(I)语言无法表达的结构;(II)从训练分布中统计不可识别的函数;(III)可表示但基于梯度训练无法达到的函数。我们为每种类型提供了一个决定性的探针。对于类型I,语言压缩残差将仅看到信号有损文本编码的文本学习者与直接看到底层信号的全信号学习者进行比较。对于类型II,反事实区分测试在与两个不兼容规则完全一致的数据上训练模型。对于类型III,盆地逃逸映射展示了一个可100%表示但标准训练达到率为0%的函数。我们发布了探针套件并讨论了对基准设计、能力审计和阴影感知不确定性的影响。

英文摘要

Some limits on what language models know are not gaps in data coverage but structural properties of learning from text. We introduce the information shadow: the region of phenomena that a text-trained learner cannot acquire regardless of scale, comprising (I) structures language cannot express, (II) functions that are statistically non-identifiable from the training distribution, and (III) functions that are representable but unreachable by gradient-based training. We give each type a probe that is decisive because the premise of the shadow is, in that setting, provable. For Type I, Language Compression Residuals compare a text learner, which sees only a lossy text-like encoding of the signal, against a full-signal learner, which sees the underlying signal directly. The text learner sits at a computable expressibility ceiling while the full-signal learner pulls away by a gap that stays flat across 300x more data, so the deficit is a property of the channel, not of training. For Type II, the Counterfactual Distinction Test trains models on data exactly consistent with two incompatible rules. Across a provable string task and a language-like agreement task, behavior on counterfactuals is set by the model's inductive bias, while 5% disambiguating data steers the learned rule bidirectionally to either target (r = +/-1.0, p < 1e-10). For Type III, Basin Escape Mapping exhibits a function that is representable at 100% (by hand construction) yet reached 0% of the time by standard training and instantly from a nearby initialization, with width scaling providing no help (p = 1.6 x 10^-14). Each effect is isolated by a control that rules out a capacity or modality artifact. We release the probe suite and discuss implications for benchmark design, capability auditing, and shadow-aware uncertainty.

↑