arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于带势函数的自回归语言模型估计期望

Estimating great expectations under autoregressive language models with potentials

Francesco I. Re, Shubhangi Ghosh, Tim Vieira, Ryan Cotterell

arXiv 2610.11399首次发表:更新:

发表机构

ETH Zürich; Columbia University(苏黎世联邦理工学院; 哥伦比亚大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出利用采样副产品的下一词元条件概率,通过可加分解测试函数的前缀势函数构造估计量,推导其降方差条件,开发实用势函数并在相当计算成本下实现显著方差降低,提升了自回归语言模型下期望估计的效率。

AI 中文摘要

语言模型的诸多应用并非依赖单个样本,而是依赖模型下测试函数的期望,可靠估计此类期望的计算成本可能很高。本文展示如何利用采样过程中作为副产品的下一个词元条件概率来提升估计效率,具体通过势函数实现:势函数是前缀上的实值函数,可将测试函数加性分解。我们构造了一个方差取决于所选势函数的估计量,并推导了势函数可降低方差的条件。随后,我们为多个待估计量和应用开发了实用势函数,在相当的计算成本下,针对多个待估计量展示了显著的方差降低效果。

英文摘要

Many applications of language models hinge not on individual samples but on the expectation of a test functional under the model. Estimating such expectations reliably can be computationally expensive. In this paper, we show how to make estimation more efficient by exploiting the next-token conditional probabilities which are available as a by-product of sampling. We do so through potentials: real-valued functions on prefixes that decompose the test functional additively. We construct an estimator whose variance depends on the chosen potential, and derive conditions under which a potential reduces this variance. We then develop practical potentials for several estimands and applications, and demonstrate substantial variance reductions across several estimands at comparable computational cost.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑