arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.16001cs.CL

BayesPO:通过并行回火梯度引导离散MCMC进行贝叶斯提示优化

BayesPO: Bayesian Prompt Optimization via Parallel-Tempered Gradient-Guided Discrete MCMC

Junjie Zhou, Zhijian Ou

中文总结 AI 辅助

研究将提示优化作为离散提示令牌上的贝叶斯后验采样问题,提出BayesPO框架,结合任务似然项与语言模型先验定义后验分布,用马尔可夫链蒙特卡罗实例化,实验表明其能发现有意义提示、逃离局部最优并提高准确率,同时揭示了两个主要局限。

中文摘要 AI 辅助

提示优化在不更新模型参数的情况下调整大语言模型,但许多自动提示优化器仍是对候选指令的启发式搜索过程。本文将提示优化作为离散提示令牌上的贝叶斯后验采样进行研究。通过结合奖励解释输入输出示例的任务似然项和支持流畅指令的语言模型先验来定义后验分布,将其转化为基于能量的后验采样问题,可用梯度引导离散MCMC提议。文中框架BayesPO用马尔可夫链蒙特卡罗实例化,采用Metropolis-Hastings校正的Gibbs-with-Langevin提议并集成并行回火。实验表明该采样器能在诊断任务中发现语义有意义的提示,并行回火有助于在诗歌完成任务中逃离局部最优,对24个指令归纳子任务的APE提示进行后优化可提高平均准确率。研究还揭示了能量最小化可能过拟合小优化集以及当前采样器计算成本高这两个主要局限。这些发现将贝叶斯提示采样定位为有原则的后优化工具,并指出了概率提示优化的一个有前景的方向。

英文摘要

Prompt optimization adapts large language models (LLMs) without updating model parameters, but many automatic prompt optimizers remain heuristic search procedures over candidate instructions. This paper studies prompt optimization as Bayesian posterior sampling over discrete prompt tokens. We define a posterior distribution by combining a task likelihood term, which rewards prompts that explain input-output examples, with a language-model prior, which favors fluent instructions. This converts prompt optimization into an energy-based posterior sampling problem, for which gradients can be used to guide discrete Markov chain Monte Carlo (MCMC) proposals over vocabulary tokens. We refer to our framework as BayesPO, short for Bayesian Prompt Optimization. In this paper, BayesPO is instantiated with Markov chain Monte Carlo: it uses a Metropolis-Hastings corrected Gibbs-with-Langevin (GwL) proposal and integrates parallel tempering for global exploration of rugged LLM-induced energy landscapes. The concrete sampler further adapts the GwL sampler to the practical constraints of non-weight-tied LLM embeddings. Experiments with Qwen2.5 models show that the sampler discovers semantically meaningful prompts on diagnostic tasks, that parallel tempering helps escape a local optimum in a poetry completion task, and that post-optimizing APE prompts on 24 instruction-induction subtasks improves average accuracy from 60.04% to 63.23%. The study also reveals two main limitations: energy minimization may overfit small optimization sets, and the current sampler remains computationally expensive. These findings position Bayesian prompt sampling as a principled post-optimization tool and point to a promising direction for probabilistic prompt optimization.

↑