arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

超越重复采样:学习用于LLM推理的搜索策略

Beyond Repeated Sampling: Learning Search Policies for LLM Reasoning

Ismail Labiad, Matthieu Kowalski, Marc Schoenauer, Rémi Munos, Julia Kempe

arXiv 2609.26704首次发表:更新:

发表机构

Meta FAIR; Université Paris-Saclay, LISN, Inria, CNRS; NYU Courant Institute and CDS(Meta FAIR; 巴黎-萨克雷大学,LISN,法国国家信息与自动化研究所,法国国家科学研究中心; 纽约大学库朗数学科学研究所和数据科学中心)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对LLM推理中朴素重复采样探索不足的问题,提出学习可训练的语义级概念生成搜索策略,以强化学习优化小型概念生成器,提升大型冻结答案生成器在困难数学推理上的pass@k,并具有跨模型迁移性。

AI 中文摘要

大型语言模型通过增加测试时计算量来应对困难的推理问题,但主导策略仍然是朴素的重复采样:抽取许多独立的解并希望其中一个正确。由于这种采样仅通过局部解码噪声进行探索,它往往产生许多近似重复的尝试,而非真正不同的想法。我们提出疑问:能否在语义层面引导探索,即首先采样问题特定的概念、提示或策略,然后基于它们生成答案。我们将此提炼为一个更简单的、更具探索性的过程,该过程在单个轨迹中生成许多多样的概念,并在重复采样表现不佳的难题上对其进行评估。我们进一步使概念生成可训练:通过强化学习优化一个小型概念生成器,使其概念最大化一个更大的、冻结的答案生成器的下游成功率。在困难的数学推理问题上,训练后的概念生成器在相同的答案生成分配下,相较于朴素的重复采样,显著提升了答案生成器的pass@k,超越了从更大的未调优模型中抽取的概念,并迁移到它从未训练过的答案生成器,包括来自不同家族的模型。因此,一个小模型可以被训练成一个有效的、可复用的搜索策略,用于一个更大的模型。

英文摘要

Large language models increasingly tackle hard reasoning problems by spending more test-time compute, yet the dominant strategy remains naive repeated sampling: draw many independent solutions and hope one is correct. Because such sampling explores only through local decoding noise, it tends to produce many near duplicate attempts rather than genuinely different ideas. We ask whether exploration can instead be steered at a semantic level, by first sampling problem specific concepts, hints, or strategies and then conditioning answer generation on them. We refine this into a simple, more exploratory procedure that emits many diverse concepts in a single trajectory, and evaluate it on hard problems where repeated sampling struggles. We then go a step further and make concept generation trainable: a small concept generator is optimized with reinforcement learning so that its concepts maximize the downstream success of a larger, frozen answer generator. On hard mathematical reasoning problems, the trained concept generator substantially improves the answer generator's pass@k over naive repeated sampling at the same answer generation allocation, surpasses concepts drawn from much larger untuned models, and transfers to answer generators it was never trained against, including a model from a different family. A small model can thus be trained into an effective, reusable search policy for a much larger one.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑