压缩的代价:事实性幻觉的率失真极限
The Cost of Compression: A Rate-Distortion Limit on Factual Hallucination
- Hefei Institutes of Physical Science, Chinese Academy of Sciences(中国科学院合肥物质科学研究院)
- University of Science and Technology of China(中国科学技术大学)
- Chongqing University(重庆大学)
- Shenzhen University(深圳大学)
机构由 AI 辅助整理,请以论文原文为准。
中文总结 AI 辅助
本文提出覆盖-压缩模型,证明闭卷问答中事实性幻觉存在率失真极限,将误差分解为压缩失真与覆盖缺失,为理解有限记忆下的有损回忆提供信息论框架。
中文摘要 AI 辅助
闭卷问答中的事实性幻觉常被视为覆盖问题:模型失败是因为相关事实不在其内部记忆中。这种观点忽略了错误的第二个来源。即使一个事实已被观察到,有限的记忆也可能迫使其仅被近似存储。我们通过一个简单的覆盖-压缩事实回忆模型研究这一效应。我们考虑一个非结构化的问答任务,有$N$个可能的查询和$K$个可能的答案。学习者观察$M$个训练事实,将它们压缩至最多$B$比特,并在无检索的情况下回答均匀抽取的测试查询。对于均匀随机的真实映射,我们证明$\mathcal{E} \geq \frac{M}{N}\delta^\star\\!\left(\frac{B}{M}\right) + \left(1-\frac{M}{N}\right)\left(1-\frac{1}{K}\right)$,其中$\delta^\star(r)$是在零一损失下均匀$K$元信源的逆率失真函数。这两项分别将观察事实的压缩失真与未观察事实的覆盖缺失分开。该界限为推理选择性记忆、强制压缩、结构、检索、弃权(不执行)和长上下文组织提供了一种紧凑的方式。我们通过理论隐含的模拟和受控事实注入探针在现代语言模型中研究预测的特征,这些探针改变事实负载和有效可训练记忆。该结果并非幻觉的完整理论,而是对一种可分离故障模式的信息论解释:在有限记忆下对观察事实的有损回忆。
英文摘要
Factual hallucination in closed-book question answering is often treated as a coverage problem: a model fails because the relevant fact is absent from its internal memory. This view misses a second source of error. Even when a fact has been observed, finite memory may force it to be stored only approximately. We study this effect through a simple coverage--compression model of factual recall. We consider an unstructured question-answering task with $N$ possible queries and $K$ possible answers. A learner observes $M$ training facts, compresses them into at most $B$ bits, and answers uniformly drawn test queries without retrieval. For a uniformly random ground-truth mapping, we prove $\mathcal{E} \geq \frac{M}{N}δ^\star\!\left(\frac{B}{M}\right) + \left(1-\frac{M}{N}\right)\left(1-\frac{1}{K}\right)$, where $δ^\star(r)$ is the inverse rate-distortion function of a uniform $K$-ary source under zero-one loss. The two terms separate compression distortion on observed facts from missing coverage on unobserved facts. The bound gives a compact way to reason about selective memory, forced compression, structure, retrieval, abstention, and long-context organization. We study the predicted signatures with theory-implied simulations and controlled fact-injection probes in modern language models that vary fact load and effective trainable memory. The result is not a complete theory of hallucination, but an information-theoretic account of a separable failure mode: lossy recall of observed facts under finite memory.