发表机构
Indian Institute of Management Bangalore(班加罗尔印度管理学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究语言模型可靠性和扩展性的信息论极限,通过分析任务输出不确定性等因素得出缩放定律,指出LLM性能受训练数据或模型容量限制,还统一多种实际现象,提供生成语言模型性能极限的统一理论。
AI 中文摘要
大语言模型(LLMs)被评估为在足够规模下对任何任务都能实现完美可靠性。但我们表明此假设在信息论上不合理。每个生成任务都有可靠性上限,由可从可观察上下文解决的输出不确定性决定。差距分为可通过额外上下文缩小的可解决部分和任务模糊性固有的主观部分。自回归生成以任务依赖内核规定的速率进一步降低此上限。由此得出第一性原理缩放定律,LLM性能受稀缺资源(训练数据或模型容量)瓶颈限制。该定律将Chinchilla缩放定律作为特殊情况涵盖,并解释了缩放何时提高可靠性。此外,我们的框架统一了各种实际现象,如检索增强的好处和灾难性遗忘的频谱机制。我们的工作形式化了跨领域模型性能的资源复杂性权衡,为生成语言模型的性能极限提供了统一理论。
英文摘要
Large language models (LLMs) are trained and evaluated as though perfect reliability is achievable for any task given sufficient scale. We show that this assumption is information-theoretically unjustified. Every generative task has a reliability ceiling that no model can exceed, determined by how much output uncertainty is resolvable from observable context. The gap decomposes into a resolvable component closable with additional context and a subjective component inherent to task ambiguity. Autoregressive generation further degrades this ceiling at a rate governed by the task's dependency kernel, which quantifies inter-token correlations in the output. From these two primitives, we derive a first-principles scaling law where LLM performance is bottlenecked by the scarcer resource: training data or model capacity. This law recovers the Chinchilla scaling law as a special case and provides a structural account of when scaling improves reliability. Beyond scaling, our framework unifies diverse practical phenomena, such as the benefits of retrieval-augmentation and the spectral mechanics of catastrophic forgetting. Our work formalizes the resource-complexity tradeoffs that govern model performance across domains, offering a unified theory of performance limits in generative language models.
Comments45 pages, 2 figures