奖励率策略梯度用于高效机器学习工程智能体
Reward-rate Policy Gradient for Efficient Machine Learning Engineering Agents
- New York University(纽约大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对智能体动作耗时可变的问题,提出奖励率策略梯度(RPG)优化单位时间长期奖励,理论验证并实验表明在MLE-Bench和NanoGPT上显著优于普通强化学习。
AI中文摘要:
传统强化学习(RL)技术侧重于最大化期望累积奖励,其中每个动作假设占用恒定的单位时间。然而,这一假设并不适用于智能体强化学习任务,如机器学习工程(MLE)智能体,其动作涉及数据加载、特征工程和模型训练,这些动作耗时可变。在现代智能体强化学习中,动作代价高昂,效率至关重要。为解决这一局限,我们借鉴连续时间强化学习和半马尔可夫决策过程(SMDP)的公式化方法,提出了奖励率策略梯度(RPG),其核心是优化奖励率——即单位时间内的长期奖励。RPG从离策略样本中估计奖励率,然后按该奖励率对每个动作所消耗的时间进行计费。我们首先在老虎机(bandit)设置下进行理论分析,证明RPG能够逼近最优奖励率,并通过实验证明其优于基线方法,同时避免了现有方法中已知的遍历策略空间的问题。随后,我们将RPG应用于带有自我改进循环的小型语言模型(Qwen3.5-4B),实验表明,在固定时间预算内,RPG在MLE-Bench和NanoGPT上获得的奖励高于普通强化学习,分别高出19.2%和85.7%。我们的方法为现代智能体强化学习任务中在等待时间考量下优化性能提供了实用解决方案,在这些任务中,动作与外部环境交互并消耗时间。
英文摘要:
Traditional reinforcement learning (RL) techniques focus on maximizing expected cumulative reward, where each action assumes to take a constant unit of time. However, this assumption does not hold for agentic RL tasks such as machine learning engineering (MLE) agents, where actions involve data loading, feature engineering, and model training that take variable durations. Efficiency matters in modern agentic RL where actions are costly. To address this limitation, we adapt from continuous-time RL and Semi-Markov Decision Process (SMDP) formulation and propose Reward-rate Policy Gradient (RPG), where we focus on optimizing the reward rate -- the long-term reward per unit of time. RPG estimates the reward rate from off-policy samples, then charges each action for the time it consumes at that rate. We first conduct theoretical analysis in the bandit setting to establish that RPG approximates the optimal reward rate and empirically demonstrate it outperforms baselines while avoiding enumeration over the policy space, a known issue for an existing method. We then further apply RPG on a small language model (Qwen3.5-4B) with self-improvement loops and empirically show it obtains higher rewards within a fixed time budget than vanilla RL on MLE-Bench and NanoGPT, with a 19.2% and 85.7% margin, respectively. Our method provides a practical solution for optimizing performance under wait time considerations in modern agentic RL tasks, where actions interact with external environments and cost time.