arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.02149cs.AI

超越均值:面向大语言模型推理的多时刻策略优化

Beyond the Mean: Multi-Moment Policy Optimization for LLM Reasoning

Yijun Zhang, Yule Xie, Jiaxin Ding, Xin Ding, Fan Xu, Haoxiang Zhang, Luoyi Fu

AI总结:

本文针对大语言模型推理提出多时刻策略优化(MMPO)框架,联合最小化失败概率分布的多个矩,在五个数学推理基准上优于基线方法,为策略优化目标设计提供新视角。

AI中文摘要:

强化学习已成为提升大语言模型推理能力的核心范式,现有方法通常旨在降低各类问题引发的失败概率。本文针对大语言模型推理的策略优化引入基于矩的视角,将随机采样问题的失败概率视为随机变量,通过其矩刻画优化目标。在该视角下,许多现有方法仅优化失败概率分布的单一矩,未能充分表征其更广泛的分布结构。我们提出多时刻策略优化(Multi-Moment Policy Optimization,MMPO),这是一种新颖的策略优化框架,可联合最小化失败概率分布的多个矩。MMPO可直接解释为最小化获得首个成功响应所需的期望截断时间。除MMPO外,我们还开发了一种通用的矩转换框架,该框架可系统地生成不同的矩轮廓,并为更广泛的策略优化目标族提供统一视角。在五个数学推理基准及不同规模模型上开展的实验表明,MMPO始终优于强劲的基线方法。我们希望这种基于矩的视角能为大语言模型推理的策略优化目标设计提供新见解。

英文摘要:

Reinforcement learning has become a central paradigm for improving the reasoning capabilities of large language models. Existing methods generally aim to reduce the failure probabilities induced across problems. In this paper, we introduce a moment-based perspective on policy optimization for LLM reasoning by treating the failure probability of a randomly sampled problem as a random variable and characterizing optimization objectives through its moments. Under this perspective, many existing methods optimize only a single moment of the failure-probability distribution, leaving its broader distributional structure largely uncharacterized. We propose \textbf{M}ulti-\textbf{M}oment \textbf{P}olicy \textbf{O}ptimization (MMPO), a novel policy optimization framework that jointly minimizes multiple moments of the failure-probability distribution. MMPO admits a direct operational interpretation as minimizing the expected truncated time required to obtain the first successful response. Beyond MMPO, we further develop a general moment-transformation framework that systematically induces different moment profiles and provides a unified view of a broader family of policy optimization objectives. Experiments across five mathematical reasoning benchmarks and models of different scales demonstrate that MMPO consistently outperforms strong baselines. We hope this moment-based perspective offers new insights into the design of policy optimization objectives for LLM reasoning.

↑