arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.26120cs.CLcs.LG

通过采样引导和扩展大语言模型的方法

Recipes for Steering and Scaling LLMs via Sampling

Jiajun He, Zongyu Guo, José Miguel Hernández-Lobato, Yuanqi Du

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出一种灵活的理论框架,基于序贯蒙特卡洛和副本交换两种算法,用于通过采样引导和扩展大语言模型,其扩展性优于Best-of-N和标准MCMC基线。

中文摘要 AI 辅助

大语言模型(LLMs)是概率模型,通常由自回归分解定义。尽管近期研究已开始探索超越基础模型的更丰富目标分布,但采样策略仍效率低下。本文提出一种灵活且具有理论基础的框架,用于通过采样引导和扩展自回归大语言模型。在该框架内,我们描述了两种算法:一种基于序贯蒙特卡洛(SMC),另一种基于副本交换(RE),可引导生成过程朝向基础模型分布的幂、乘积或倾斜方向。我们通过在无外部监督或奖励模型的情况下扩展大语言模型的生成质量来阐释该框架。实验结果表明,我们的方法相比Best-of-N和标准MCMC基线具有更优的扩展性。总体而言,本文为通过采样进行大语言模型的概率推理提供了一套系统的方法。

英文摘要

Large Language Models (LLMs) are probabilistic models, typically defined by an autoregressive factorization. While recent work has begun to study richer target distributions beyond the base model, the sampling strategies remain highly inefficient. In this paper, we present a flexible and theoretically grounded framework for steering and scaling autoregressive LLMs with sampling. Within this framework, we describe two algorithms -- one based on Sequential Monte Carlo (SMC) and one based on Replica Exchange (RE) -- that steer generation toward powering, product or tilting of the base model distribution. We illustrate this framework through scaling the generation quality of LLMs without external supervision or reward models. Experimental results demonstrate our methods scale more favorably than Best-of-N and standard MCMC baselines. Overall, this paper offers a systematic recipe for probabilistic inference with LLMs via sampling.

发表机构

  • Microsoft(微软公司)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑