arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

无需反向传播的启发:通过优化潜在后验来引导模型行为

Elicitation without Backpropagation: Steering Model Behavior by Optimizing the Latent Posterior

Garrett Baker, Vinayak Pathak, Daniel Murfet, Susan Wei

arXiv 2607.18804首次发表:更新:

发表机构

Monash University; Timaeus(莫纳什大学; 提马尤斯)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究利用潜在后验模型提出后验前缀调整(PPT)方法,在贝叶斯滤波变压器中,给定效用函数找高预期效用提示。通过预测蒙特卡洛抽取先验样本,用重要性采样估计梯度,无需前向传递和反向传播,一组样本可驱动针对多种效用的启发并验证。

AI 中文摘要

在变压器行为的潜在后验模型中,下一个token分布源于基于上下文的潜在预测模型的后验,并混合生成延续。我们在精确的设置中利用此模型,即从分层先验的序列中进行元学习的贝叶斯滤波变压器(BFT),引入了后验前缀调整(PPT),这是一种从变压器中“引出”行为的新方法:给定延续上的效用函数,找到一个提示,使变压器生成具有高预期效用的延续。对于BFT,启发目标通过潜在后验进行分解,并且该目标的梯度可以仅从先验样本中估计。PPT优化硬提示上分布的参数:它通过预测蒙特卡洛(PMC)从BFT中抽取一次先验样本,然后通过对它们进行重要性采样来估计梯度。该优化不执行变压器前向传递,也不通过变压器进行反向传播,并且先验样本与效用无关,因此一组样本以可忽略的边际成本驱动针对任意数量效用的启发。我们在三个效用族(反向交叉熵、频率匹配、Dyck有效性)的贝塔-伯努利和强化瓮BFT上验证了PPT。

英文摘要

In the \emph{latent posterior model} of transformer behavior, the next-token distribution arises from a posterior over latent predictive models conditioned on the context, mixed to generate continuations. We exploit this model in settings where it is exact, namely Bayes-filtered transformers (BFTs) meta-learned on sequences from a hierarchical prior, to introduce \textbf{Posterior Prefix Tuning (PPT)}, a new method for \emph{eliciting} behavior from a transformer: given a utility function on continuations, find a prompt under which the transformer generates continuations of high expected utility. For a BFT, the elicitation objective factors through the latent posterior, and the gradient of this objective can be estimated from samples of the prior alone. PPT optimizes the parameters of a distribution over hard prompts: it draws prior samples once from the BFT via predictive Monte Carlo (PMC), then estimates the gradient by importance sampling against them. The optimization performs no transformer forward passes and no backpropagation through the transformer, and the prior samples are utility-independent, so a single set of samples drives elicitation against any number of utilities at negligible marginal cost. We validate PPT on Beta--Bernoulli and reinforced urn BFTs across three utility families (reverse cross-entropy, frequency matching, Dyck validity).

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑