arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.15476cs.CLcs.LG

温度脆弱性与截断采样的条件性收益

Temperature Fragility and the Conditional Benefits of Truncation Sampling

Francesco La Rosa

首次发表
浏览论文内容

中文总结 AI 辅助

本研究系统测试了截断采样器在不同温度下对十三种大语言模型准确率的影响,发现其收益仅在高温导致性能显著下降时出现,默认温度下无优势。

中文摘要 AI 辅助

大语言模型通过从预测分布中采样每个词元来生成文本,温度参数决定了采样偏离最可能词元的程度。top-p和min-p等截断采样器在采样前丢弃概率最低的词元,从而使高温采样保持连贯性。它们报告的准确率提升来自1.5至3的温度,而部署系统的默认值集中在0.6至1.0之间。这些默认温度下截断采样器是否改变准确率,以及针对哪些模型,尚未被测量。我们在一个受控流程中,以0.7、1.0和1.3的温度对GSM8K和MMLU-Pro上的十三个开放权重模型进行了测试,其中十个模型在八种解码配置下进行了测试。十三个模型中有六个在MMLU-Pro上从0.7到1.3损失了17至38个准确率点,其余七个最多损失10个。损失的准确率来自运行到词元限制或从未陈述答案的生成结果。这些结果表明,截断采样器主要在高温度显著降低模型性能时提升准确率。在准确率跨温度保持稳定的情况下,所测试的截断采样器均未优于普通温度采样。

英文摘要

Large language models generate text by sampling each token from a predicted distribution, and a temperature parameter sets how far the draw strays from the most probable tokens. Truncation samplers such as top-p and min-p discard the least probable tokens before the draw, so that sampling at high temperature stays coherent. Their reported accuracy gains come from temperatures of 1.5 to 3, while the defaults of deployed systems cluster between 0.6 and 1.0. Whether they change accuracy at those defaults, and for which models, has not been measured. We test thirteen open-weight models on GSM8K and MMLU-Pro at temperatures 0.7, 1.0, and 1.3 in one controlled pipeline, ten of them under eight decoding configurations. Six of the thirteen models lose 17 to 38 accuracy points on MMLU-Pro between 0.7 and 1.3, and the other seven lose at most 10. The lost accuracy comes from generations that run to the token limit or never state an answer. These results suggest that truncation samplers improve accuracy primarily when higher temperatures substantially degrade model performance. Where accuracy remains stable across temperatures, none of the tested truncation samplers improves on plain temperature sampling.

发表机构

  • University of Edinburgh(爱丁堡大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑