arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.17823cs.LG

$\max$@$k$强化学习的理论基础

Theoretical Foundations of $\max$@$k$ Reinforcement Learning

Riccardo Poiani, Martino Bernasconi, Andrea Celli

首次发表
浏览论文内容

中文总结 AI 辅助

研究有限 horizon 强化学习中$\max$@$k$学习问题,指出其与标准预期回报最大化不同,证明马尔可夫策略不足,识别状态增强,表征策略性能差距,表明学习更难,给出高效算法实现最优样本复杂度率。

中文摘要 AI 辅助

强化学习是现代大型推理模型的基石技术。对于代码生成和定理证明等困难任务,通常通过生成$K$个响应而非单个响应来评估智能体,并使用诸如$\max$@$k$等考虑重试的指标来衡量性能。尽管其具有实际重要性,但此类标准下学习的理论基础仍然有限。本文对有限 horizon 强化学习中的$\max$@$k$学习问题进行了理论研究。表明优化$\max$@$k$目标与标准预期回报最大化根本不同,证明马尔可夫策略通常不足,识别出恢复最优性的紧凑状态增强,明确表征历史依赖和非历史依赖策略之间可能出现的性能差距。还表明学习$\max$@$k$最优策略在统计上比标准强化学习更难,并提供了一种实现最优样本复杂度率的高效算法。

英文摘要

Reinforcement Learning is a cornerstone technique for modern large reasoning models. Usually, for difficult tasks such as code generation and theorem proving, the agent is evaluated by generating $K$ responses rather than sampling a single response, and performance is then measured using a retry-aware metric such as $\max$@$k$. Despite their practical importance, the theoretical foundations of learning under such criteria remain limited. In this work, we provide a theoretical study of the $\max$@$k$ learning problem in finite-horizon reinforcement learning. We show that optimizing the $\max$@$k$ objectives is fundamentally different from standard expected-return maximization. In particular, we prove that Markovian policies are in general insufficient, identify a compact state augmentation that restores optimality, and explicitly characterize the performance gap that can arise between history-dependent and non-history-dependent policies. Moreover, we show that learning $\max$@$k$-optimal policies is statistically harder than standard reinforcement learning and provide an efficient algorithm that achieves the optimal sample complexity rate.

发表机构

  • Bocconi University(博科尼大学)

机构由 AI 辅助整理,请以论文原文为准。

↑