上下文学习的贝叶斯缩放定律
Bayesian scaling laws for in-context learning
- Stanford University(斯坦福大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文提出贝叶斯缩放定律解释上下文学习,证明ICL近似贝叶斯学习器,并通过GPT-2和指令微调LLM实验揭示后训练难以阻止被抑制能力在ICL下重新出现。
AI中文摘要:
上下文学习(ICL)是一种强大的技术,可在不进行训练更新的情况下让语言模型执行复杂任务。先前的工作已确立所提供的上下文示例数量与模型预测准确率之间的强相关性。在本文中,我们试图通过证明 ICL 近似于贝叶斯学习器来解释这种相关性。这一视角为 ICL 提出了一种新颖的贝叶斯缩放定律。在不同规模的 GPT-2 模型实验中,我们的缩放定律在准确率上与现有缩放定律相匹配,同时提供了可解释的项,用于刻画任务先验、学习效率和逐示例概率。为了说明这种可解释缩放定律所提供的分析能力,我们报告了受控合成数据集实验,这些实验旨在为现实世界中的安全对齐研究提供信息。在我们的实验方案中,我们使用 SFT 或 DPO 来抑制模型已有的不期望能力,然后使用 ICL 尝试恢复该能力(多示例越狱)。随后,我们使用能力基准以及一个新的多示例越狱数据集,在现实世界的指令微调 LLM 上研究 ICL。在所有情况下,贝叶斯缩放定律都能准确预测 ICL 导致被抑制行为重新出现的条件,这揭示了后训练在提高 LLM 安全性方面的无效性。
英文摘要:
In-context learning (ICL) is a powerful technique for getting language models to perform complex tasks with no training updates. Prior work has established strong correlations between the number of in-context examples provided and the accuracy of the model's predictions. In this paper, we seek to explain this correlation by showing that ICL approximates a Bayesian learner. This perspective gives rise to a novel Bayesian scaling law for ICL. In experiments with \mbox{GPT-2} models of different sizes, our scaling law matches existing scaling laws in accuracy while also offering interpretable terms for task priors, learning efficiency, and per-example probabilities. To illustrate the analytic power that such interpretable scaling laws provide, we report on controlled synthetic dataset experiments designed to inform real-world studies of safety alignment. In our experimental protocol, we use SFT or DPO to suppress an unwanted existing model capability and then use ICL to try to bring that capability back (many-shot jailbreaking). We then study ICL on real-world instruction-tuned LLMs using capabilities benchmarks as well as a new many-shot jailbreaking dataset. In all cases, Bayesian scaling laws accurately predict the conditions under which ICL will cause suppressed behaviors to reemerge, which sheds light on the ineffectiveness of post-training at increasing LLM safety.