arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

语言建模是单调压缩

Language Modeling is Monotone Compression

Noam Mazor, Andrew Morgan, Rafael Pass

arXiv 2610.11031首次发表:更新:

发表机构

New York University; Cornell Tech; Technion; Tel-Aviv University(纽约大学; 康奈尔科技; 以色列理工学院; 特拉维夫大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究从理论上证明大型语言模型等价于单调压缩算法,且该等价关系需密码学单向函数存在,还推导了下一位伪熵与单调不可压缩性等价的密码学结论。

AI 中文摘要

人工智能与神经科学领域有一个长期存在的假说,认为智能与压缩密切相关:高效压缩信息的能力直观反映了与智能和学习相关的能力。近期实验研究验证了这一直觉,表明大型语言模型(LLMs)的能力与其作为压缩器的能力存在关联:例如,Deletang等人(ICLR'24)证明LLMs可作为强大的压缩器,Huang等人(COLM'24)则发现LLMs的压缩能力与其在知识和推理基准上的表现高度相关。本研究针对这一关联展开理论研究,主要结果是:形式化为下一个词预测器的LLMs等价于单调(即保序)压缩算法——这类算法的编码过程保留输入的顺序,且二者可相互构造,误差仅存在加性间隙2。进一步证明,当且仅当存在密码学意义上的(无限频繁)单向函数时,该等价关系才需要单调性。由此得到一个具有独立意义的密码学推论:分布的下一位伪熵(熵的计算模拟概念)等价于该分布的单调不可压缩性,此前仅已知不可压缩性蕴含下一位伪熵(Haitner等人,ITCS'23)。

英文摘要

A long-standing hypothesis in artificial intelligence and neuroscience posits that intelligence is closely related to compression: the ability to compress information efficiently intuitively reflects capacities associated with intelligence and learning. Indeed, recent experimental works verify this intuition by showing connections between the capabilities of large language models (LLMs) and their ability as compressors: for instance, Deletang et al. (ICLR'24) demonstrate that LLMs can be used as powerful compressors, and Huang et al. (COLM'24) show that the compression ability of LLMs is highly correlated with their performance on benchmarks for knowledge and reasoning. In this work, we initiate a theoretical study of this connection. Our main result is that LLMs (formally modeled as next-token predictors) are equivalent to monotone (a.k.a. order-preserving) compression algorithms---namely, compression algorithms where the encoding process preserves the ordering of the inputs---in the sense that the one can be constructed from the other while preserving the same error up to an additive gap of 2. We next show that the monotonicity is required for this equivalence to hold if and only if cryptographic (infinitely-often) one-way functions exist. As a direct corollary, we get a cryptographic result of independent interest: the notion of next-bit pseudoentropy (a computational analogue of entropy) of a distribution is equivalent to monotone incompressibility of the distribution. (Previously, it was only known (Haitner et al., ITCS'23) that incompressibility implies next-bit pseudoentropy.)

Comments22 pages, 1 figure. Submitted to ICLR 2027

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑