arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

卡尔曼与课程学习结合:面向自适应RL微调的高效动态提示选择

Kalman Meets Curriculum: Efficient Dynamic Prompt Selection for Adaptive RL Finetuning

Haodong Zhu, Yangyang Ren, Yanjing Li, Sheng Xu, Haiguang Liu, Linlin Yang, Baochang Zhang

arXiv 2607.27610首次发表:更新:

发表机构

Beihang University; Zhongguancun Academy; Nanyang Technological University; Communication University of China(北京航空航天大学; 中关村学院; 南洋理工大学; 中国传媒大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出KGPS方法,将提示选择转化为动态状态估计问题,适配RL训练动态,在多推理基准和RL算法上较基线提升准确率与rollout效率,在在线提示选择中达最优。

AI 中文摘要

强化学习(RL)微调可显著提升大语言模型(LLM)的推理能力,但其效果高度依赖于为当前策略选择难度合适的提示,而提示难度会随训练过程变化,这一问题极具挑战性。现有在线方法面临权衡:基于评估的方法准确但成本高昂,基于预测的方法高效但通常假设难度平稳,不适合RL的非平稳训练动态。为解决这些问题,本文提出卡尔曼引导提示选择方法(KGPS),将提示选择重新表述为动态状态估计问题而非静态难度预测。KGPS在logit空间中使用线性高斯状态空间模型对每个提示的潜在成功率进行建模,过程噪声与策略更新幅度耦合,使得策略变化更显著时不确定性增加。卡尔曼滤波器随后维持提示难度的校准高斯后验,通过最大化后验期望训练效用选择提示,该效用偏好中等难度提示并自然重新访问不确定提示。该过程对策略漂移具有适应性,且无需额外的rollout,仅需标准策略训练。在数学、规划和几何推理基准及多种RL算法上开展的大量实验表明,KGPS相较于强基线始终提升最终准确率和rollout效率,在在线提示选择方法中达到当前最优性能。例如,在DeepSeek-R1-Distill-7B上,KGPS相较于DS减少83%的rollout,同时在六个数学推理基准上平均性能提升0.12个点。

英文摘要

Reinforcement learning (RL) finetuning significantly enhances the reasoning capabilities of large language models (LLMs), yet its effectiveness critically depends on selecting prompts of appropriate difficulty for the current policy. This is challenging because prompt difficulty evolves throughout training. Existing online methods therefore face a trade-off: evaluation-based approaches are accurate but expensive, while prediction-based approaches are efficient but typically assume stationary difficulty, making them ill-suited to RL's non-stationary training dynamics. To address these issues, we propose a Kalman-Guided Prompt Selection method (KGPS), which reformulates prompt selection as a dynamic state estimation problem rather than static difficulty prediction. KGPS models each prompt's latent success rate in logit space using a linear-Gaussian state-space model, with process noise coupled to the magnitude of policy updates so that uncertainty increases when the policy changes more substantially. A Kalman filter then maintains a calibrated Gaussian posterior over prompt difficulty, and prompts are selected by maximizing a posterior-expected training utility that favors intermediate-difficulty prompts while naturally revisiting uncertain ones. The resulting procedure is adaptive to policy drift and requires no additional rollouts beyond standard policy training. Extensive experiments across mathematics, planning, and geometry reasoning benchmarks, as well as multiple RL algorithms, show that KGPS consistently improves both final accuracy and rollout efficiency over strong baselines, establishing state-of-the-art performance among online prompt selection methods. For example, on DeepSeek-R1-Distill-7B, KGPS uses 83% fewer rollouts than DS while even improving the average performance by 0.12 point across six math reasoning benchmarks.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑