规划学习
Planning to Learn
浏览论文内容
中文总结 AI 辅助
针对策略梯度在分类中的短视问题,提出视界损失,通过截断交叉熵的总误差,在训练中从交叉熵过渡到精确策略梯度,在MNIST和ImageNet上提升了准确率。
中文摘要 AI 辅助
策略梯度方法是现代强化学习的核心,包括大语言模型的后训练。当它们遇到困难时,常见的嫌疑是探索、信用分配和动作采样噪声。分类问题则没有这些问题。分类器是一个策略,其期望奖励,即其“期望准确率”,是它分配给正确标签的概率,并且由于该标签是已知的,策略梯度是精确且平滑的。然而,即使在期望准确率上,精确的策略梯度也输给了交叉熵。精确梯度是短视的:它只根据当前能买到的东西来评价一次更新,但每次更新也决定了下一次更新的起点,因此一次更新的价值取决于还剩下多少学习空间。从这个角度看,交叉熵是耐心的准确率,即一个样本如果其对数几率以单位速度永远上升所支付的总误差,而精确策略梯度是零视界极限。将这一总量截断在剩余的学习量处,就得到了视界损失,这是一个一行代码的改动,随着训练进行,从交叉熵向精确策略梯度移动。在一个简单的分配模型中,它被证明能逃脱每个端点所陷入的陷阱。在MNIST和ImageNet上,使用ResNet-50、ResNet-101和ViT-S/16,视界损失在平坦学习率下比交叉熵提高了top-1准确率,并且随着标签噪声的增加,增益也随之增大。
英文摘要
Policy-gradient methods are central to modern reinforcement learning, including LLM post-training. When they struggle, the usual suspects are exploration, credit assignment and action-sampling noise. Classification has none of them. A classifier is a policy whose expected reward, its \emph{expected accuracy}, is the probability it assigns to the correct label, and because that label is known, the policy gradient is exact and smooth. Yet exact policy gradient loses to cross-entropy, even on expected accuracy. The exact gradient is myopic: it values an update only by what it buys now, but each update also sets where the next one starts, so an update's value depends on how much learning remains. Viewed this way, cross-entropy is patient accuracy, the total error an example would pay if its log-odds rose at unit speed forever, while exact policy gradient is the zero-horizon limit. Truncating this total at the learning that remains yields the horizon loss, a one-line change that moves from cross-entropy toward exact policy gradient as training runs out. In a simple allocation model, it provably escapes the trap that catches each endpoint. On MNIST and on ImageNet with ResNet-50, ResNet-101 and ViT-S/16, the horizon loss improves top-1 accuracy over cross-entropy at a flat learning rate, and the gain grows with label noise.
发表机构
- Google DeepMind(谷歌DeepMind)
机构由 AI 辅助整理,请以论文原文为准。