杠杆学习:每接收一比特信息所摧毁的熵
Leveraged Learning: entropy cleared per bit received
浏览论文内容
中文总结 AI 辅助
本文提出杠杆概念(每接收一比特信息所摧毁的熵),证明智能先验可提高杠杆,并构造简单性先验使任意非递增轮廓在宏观时间可实现。
中文摘要 AI 辅助
一个学习者在布尔映射上持有先验信念,这些映射回答一个包含 $Q$ 个问题的有限集合,并且学习者逐个接收答案。每个答案都消耗惊奇度并摧毁不确定性,不仅关于已提出的问题,还关于所有尚未提出的问题。我们将这个比率称为杠杆:每接收一比特惊奇度所摧毁的表熵。对于均匀先验,该比率为单位值,但在智能先验下可以更高(不会更低)。对真实先验和问题顺序取平均后,杠杆被精确求出,并由一个序列生成:$\ell$ 个问题的答案的平均熵 $G_\ell$。在固定提问比例 $t = \ell/Q$ 下,将输入比特数推向无穷大,该序列的增量成为一个轮廓 $\gamma(t)$,以及初始问题熵 $\eta_0$。杠杆趋近一个热力学极限。$L(t) = [\eta_0 - (1-t)\gamma(t)]/\int_0^t \gamma$。根据 de Finetti 定理,可交换先验都给出平坦的 $\gamma(t)$,因此得到双曲型的 $L(t)$,它们的演绎被限制在 $t = 0$ 处的边界层内。我们构造了一个简单性先验来摆脱这一限制,通过布尔映射在 $\mathbb{F}_2$ 上的多项式次数对其进行分级,并通过一个累积分布函数 $F$ 在次数壳层间分配权重。Reed-Muller 容量随后精确给出 $\gamma(t) = 1 - F(t)$,因此任何非递增轮廓及其生成的任何杠杆曲线,都能在宏观时间尺度上实现。
英文摘要
An expository essay containing some new results. A learner holds a prior belief over Boolean maps that answer a finite set of $Q$ questions, and receives answers one by one. Each answer carries surprisal and can clear predictive uncertainty about both the question asked and questions still unasked. We quantify this effect by the leverage: table entropy cleared per bit of surprisal received. Taking expectations over the truth prior and the question order, define the aggregate leverage as the ratio of expected uncertainty cleared to expected surprisal received. It equals one for independent answers and can exceed one for correlated answers. At finite size, this ratio is determined exactly by a single sequence: the mean entropy $G_\ell$ of the answers to $\ell$ questions. As the number of input bits grows at fixed asked fraction $t=\ell/Q$, a limiting increment profile $γ(t)$ determines the macroscopic learning curve. With $η_0$ the limiting initial entropy per question, the leverage becomes $L(t)= \frac{η_0-(1-t)γ(t)} {\int_0^tγ(x),dx}$. For exchangeable priors, de Finetti's representation gives a constant bulk profile $γ(t)$: deduction is confined to a boundary layer at $t=0$, and the leverage is forced to a hyperbolic form, as surprisal grows linearly. By contrast, we construct a simplicity prior with nontrivial bulk learning by grading Boolean maps by their polynomial degree over $\mathbb{F}_2$ and allocating weight across degree classes through a CDF $F$. Reed-Muller capacity then yields $γ(t)=1-F(t)$. This realizes any nonincreasing profile taking values in $[0,1]$, together with the corresponding macroscopic leverage curve.