arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.19611cs.CLcs.AIcs.LG

快速分叉:高效估计文本生成中的不确定性动态

Forking Fast: Efficiently Estimating Uncertainty Dynamics in Text Generation

Eric Bigelow, Amir Zur, Satchel Grant, Tal Haklay, Can Rager, Owen Lewis, Thomas McGrath, Jack Merullo, Ekdeep Singh Lubana, Atticus Geiger

首次发表
浏览论文内容

中文总结 AI 辅助

针对文本生成中不确定性动态估计成本高的问题,本研究开发统计模型平滑低采样推理数据以近似高采样数据,提升重采样分析效率,揭示不确定性动态的稳定模式及噪声来源。

中文摘要 AI 辅助

大型语言模型(LLM)的推理具有随机性,因此要理解该模型,需掌握其针对给定问题可能产生的推理链分布,即其不确定性。基于重采样的分析可表征该分布,揭示推理链的哪一步决定模型如何得出答案。然而,这些方法的主要局限在于,对推理链中每个token或句子的文本序列进行重采样成本极高。本研究致力于提升重采样分析的计算效率,同时阐明一个重要科学问题:解释文本生成中不确定性动态的合适统计模型是什么?我们发现,当对大量推理链进行重采样时,不确定性动态会收敛到稳定模式,噪声主要是采样的产物,而非LLM对每个单独token或推理步骤的敏感性。我们开发了一种统计模型,用于对低采样量推理数据中的噪声进行平滑处理,以更好地近似高采样量数据,从而大幅降低采样成本。

英文摘要

LLM reasoning is stochastic, and so understanding a model requires grappling with the distribution of reasoning chains that it might produce for a given question, i.e., its uncertainty. Resampling-based analyses characterize this distribution, revealing which steps of a rollout determine how the model arrives at its answer. However, a major limitation of these approaches is that resampling text sequences at every token or sentence in a reasoning chain is very costly. Our work strives to make resampling analysis more computationally efficient, while also shedding light on an important scientific question: what is the right statistical model for explaining uncertainty dynamics in text generation? We show that when resampling many reasoning chains, uncertainty dynamics converge to stable patterns, and noise is largely an artifact of sampling rather than an LLM's sensitivity to each individual token or reasoning step. We develop a statistical model for smoothing noisy low-sample rollout data to better approximate high-sample data, allowing us to significantly cut sampling costs.

↑