arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

自适应功率采样用于大语言模型推理

Adaptive Power Sampling for LLM Reasoning

Bingnan Xiao, Chenhao Yang, Bingcong Li, Wei Ni, Xin Wang

arXiv 2610.08563首次发表:更新:

发表机构

Fudan University; University of Glasgow; ETH Zurich; Edith Cowan University(复旦大学; 格拉斯哥大学; 苏黎世联邦理工学院; 埃迪斯科文大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出自适应功率采样(APS),根据查询难度动态调整锐化指数,利用自我奖励差距提升大语言模型推理性能,在多个基准上无需训练即超越固定指数方法。

AI 中文摘要

序列级功率采样最近作为一种无需训练的方法出现,通过从基础大语言模型(LLM)的锐化输出分布中进行采样来进行推理。然而,现有方法通常对所有查询统一锐化基础模型分布,忽略了查询难度的变化以及基础模型对每个查询的已有处理能力。本工作的目标是为功率采样赋予查询自适应性。理论上,我们证明了进一步锐化的收益由正确与不正确响应之间的自我奖励差距决定。基于这一见解,我们提出了自适应功率采样(APS),它在测试时利用答案一致性与模型自我奖励之间的关系,逐查询调整锐化指数。在包括MATH500、HumanEval和GPQA在内的多种推理任务上的实验表明,APS在无需额外训练的情况下,始终优于使用固定锐化指数的功率采样。

英文摘要

Sequence-level power sampling has recently emerged as a training-free approach to reasoning by sampling from a sharpened output distribution of a base large language model (LLM). Nevertheless, existing methods typically sharpen the base model distribution uniformly across queries, overlooking variations in query difficulty and in how well the base model already handles each query. The goal of this work is to equip power sampling with query adaptivity. Theoretically, we show that the benefits of further sharpening are determined by the self-reward gap between correct and incorrect responses. Based on this insight, we propose \emph{Adaptive Power Sampling} (APS), which adjusts the sharpening exponent on a per-query basis at test time using the relationship between answer agreement and the model's self-reward. Experiments across diverse reasoning tasks, including MATH500, HumanEval, and GPQA, show that APS consistently outperforms power sampling with a fixed sharpening exponent, without additional training.

Comments22 pages, 6 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑