发表机构
Massachusetts Institute of Technology; New York University; University of Tübingen(麻省理工学院; 纽约大学; 图宾根大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出贝叶斯微调方法,使语言模型在航班推荐任务中具备贝叶斯信念与推理能力,优于标准监督微调,并验证了信念编码与交换的有效性。
AI 中文摘要
语言模型(LMs)越来越多地被用于需要从少量观测中推理隐藏变量的任务,而贝叶斯推理是这类任务的规范性正确解决方案。虽然在一个最优的贝叶斯模型的输出上对语言模型进行监督微调会导致接近贝叶斯的行为,但在任务真实答案上的标准监督微调(SFT)却达不到这一水平。但仅凭行为并不能告诉我们,为什么在贝叶斯信号或oracle(真实答案)信号上微调会产生差异:结果语言模型是代表了贝叶斯信念、对其采取行动,还是像贝叶斯规则那样将其转化为选择。为了比较它们,我们为语言模型成为贝叶斯决策者制定了越来越严格的要求,涵盖其行为、表征和计算,并在一个航班推荐任务上进行了测试。贝叶斯训练的语言模型表现出贝叶斯行为,在其中间层编码了贝叶斯规则的数量,并在一定程度上使用编码的信念进行推荐。oracle训练的语言模型在持有的信念以及是否将信念读出为推荐方面都与前者不同。在语言模型之间交换信念转移了部分贝叶斯优势。因此,贝叶斯微调以一种标准SFT在oracle答案上无法做到的方式,在语言模型中安装了可用的贝叶斯信念,用于不确定性下的推理,凸显了细致监督的优势。
英文摘要
Language models (LMs) are increasingly used for tasks that require reasoning about hidden variables from a few observations, for which Bayesian inference is the normatively correct solution. While supervised fine-tuning of an LM on the outputs of an optimal $\textit{Bayesian}$ model leads to near-Bayesian behavior, standard supervised fine-tuning (SFT) on the true answers for the task falls short of it. But behavior alone does not tell us $\textit{why}$ tuning on a $\textit{Bayesian}$ or an $\textit{oracle}$ (true answers) signal differs: whether the resulting LM represents Bayesian beliefs, acts on them, or turns them into a choice the way Bayes' rule does. To compare them, we formulate increasingly demanding requirements for an LM to count as a Bayesian decision maker, spanning its behavior, representations, and computations, and test them on a flight recommendation task. The Bayes-trained LM acts Bayesian, encodes quantities of Bayes' rule in its middle layers, and uses the encoded belief for the recommendation to a certain extent. The oracle-trained LM differs from it both in the beliefs it holds and whether it reads beliefs out into recommendations. Exchanging beliefs between the LMs transfers a part of the Bayesian advantage. Bayes fine-tuning thus installs usable Bayesian beliefs in an LM for reasoning under uncertainty in a way standard SFT on oracle answers cannot, highlighting the advantage of nuanced supervision.
Commentsunder review, 27 pages, 25 figures