arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.02828cs.AIcs.CLcs.LG

FSPO:预算受限的LLM强化学习后训练中的策略一致风险与帕累托可行控制

FSPO: Policy-Consistent Risk and Pareto-Feasible Control for Budgeted LLM RL Post-Training

Miaobo Hu, Shuhao Hu, Xiaobo Guo, Xin Wang, Bokun Wang, Daren Zha, Jun Xiao

首次发表
浏览论文内容

中文总结 AI 辅助

FSPO通过策略一致风险模型、决策条件轨迹校准和帕累托资源延续证书,联合解决预算受限LLM强化学习后训练中的风险估计、校准和资源可行性问题,显著提升准确率并降低决策错误。

中文摘要 AI 辅助

自适应LLM强化学习后训练会在线调整多个训练执行器,包括rollout温度、组大小、裁剪、KL正则化、验证器分配和更新预算。三个耦合问题仍未解决。从行为轨迹训练的未来风险模型不一定能估计将要部署的控制器所引发的风险;在选择性动作选择后,基于记录的state-action对校准的分数可能变得失准;独立的每资源最小成本通常不能保证可行的多资源延续。我们引入了FSPO,一种用于预算受限的LLM RL后训练的反馈状态控制器,联合解决这些问题。FSPO学习一个策略一致的风险到目标模型,其Bellman目标遵循用于未来决策的相同冻结控制器,并配有一个长视野效用模型。决策条件轨迹校准(DCTC)在由临时控制器选择的动作生成的交叉拟合轨迹上校准风险。帕累托资源延续证书(PRCC)仅在非支配累积预留量在剩余时间范围内保持可行时才允许动作。在匹配的GRPO资源包络下,FSPO达到66.11%的保留集准确率和59.43%的分布外准确率,而最强的评估自适应基线PB2分别为64.47%和57.03%。三个配对训练种子在保留集和分布外评估上比上下文bandit分别提高了+2.42和+3.19个百分点。在高行为-部署不匹配下,策略一致风险将选择决策的ECE从0.108降至0.053;DCTC在匹配接受率下将其从0.039降至0.022;PRCC在18动作目录上消除了虚假可行接受($0.197\ ightarrow0.000$);在因子消融中,启用所有三个组件将轨迹失败率从0.181降至0.083。

英文摘要

Adaptive LLM reinforcement-learning post-training changes multiple training actuators online, including rollout temperature, group size, clipping, KL regularization, verifier allocation, and update budget. Three coupled issues remain unresolved. A future-risk model trained from behavior trajectories need not estimate the risk induced by the controller that will be deployed; a score calibrated on logged state-action pairs can become miscalibrated after selective action choice; and independent per-resource minimum costs do not in general certify a feasible multi-resource continuation. We introduce FSPO, a feedback-state controller for budgeted LLM RL post-training that addresses these issues jointly. FSPO learns a policy-consistent risk-to-go model whose Bellman target follows the same frozen controller used for future decisions, together with a long-horizon utility model. Decision-conditioned trajectory calibration (DCTC) calibrates risk on cross-fitted trajectories generated by actions selected by provisional controllers. A Pareto resource continuation certificate (PRCC) admits an action only when a non-dominated cumulative reservation remains feasible over the residual horizon. Under a matched GRPO resource envelope, FSPO reaches 66.11% held-out and 59.43% OOD accuracy, compared with 64.47% and 57.03% for PB2, the strongest evaluated adaptive baseline. Three paired training seeds give gains of +2.42 and +3.19 percentage points over the contextual bandit on held-out and OOD evaluation. Under high behavior-deployment mismatch, policy-consistent risk lowers selected-decision ECE from 0.108 to 0.053; DCTC lowers it from 0.039 to 0.022 at matched acceptance; PRCC removes false-feasible admissions on an 18-action catalog ($0.197\rightarrow0.000$); and enabling all three components reduces trajectory failure from 0.181 to 0.083 in a factorial ablation.

发表机构

  • University of Chinese Academy of Sciences(中国科学院大学)
  • Institute of Information Engineering, Chinese Academy of Sciences(中国科学院信息工程研究所)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑