SmoothRL:异步执行期间的在线强化学习
SmoothRL: Online Reinforcement Learning During Asynchronous Execution
- Astribot Team(Astribot团队)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
SmoothRL是一种在异步推理循环内微调预训练策略的在线RL框架,通过显式建模异步推理过程优化策略,在高精度及高度动态机器人任务上实现有效评估。
AI中文摘要:
将机器人策略部署到物理世界需要满足两个基本要求:可靠性和流畅的实时执行。然而,部署最先进的通用模型在这两方面都存在挑战。实现现实世界部署所需的精度和鲁棒性需要样本高效的在线强化学习(RL)来调整预训练模型。同时,机器人基础模型规模的不断扩大导致推理延迟升高。为了在高延迟下满足实时约束,现代系统采用带动作分块的异步推理,将策略计算与分块执行重叠以隐藏延迟并实现流畅控制。尽管异步推理与动作分块具有互补作用,但将异步执行与基于梯度的在线RL相结合的研究仍不充分。我们提出SmoothRL,这是一种在异步推理循环内微调预训练策略的在线RL框架。SmoothRL遵循值梯度范式,直接使用动作值函数关于策略动作的梯度更新策略参数。为了在异步执行下实现正确优化,SmoothRL在训练过程中显式建模异步推理过程。具体而言,每个生成的动作分块按帧索引划分为三个区域:提交区域,包含上一推理周期已提交的动作;执行区域,包含机器人执行的新生成动作;以及丢弃区域,包含被下一推理周期取代的动作。梯度仅通过执行区域传播,确保策略优化与异步执行诱导的轨迹分布一致。我们在需要高精度的现实机器人任务以及需要异步执行的高度动态任务上评估SmoothRL。
英文摘要:
Deploying robot policies in the physical world requires satisfying two fundamental desiderata: reliability and smooth real-time execution. However, deploying state-of-the-art generalist models presents challenges on both fronts. Achieving the precision and robustness required for real-world deployment necessitates sample-efficient online reinforcement learning (RL) to adapt pretrained models. Meanwhile, the increasing scale of robot foundation models has led to higher inference latency. To satisfy real-time constraints under high latency, modern systems adopt asynchronous inference with action chunking, overlapping policy computation with chunk execution to hide latency and enable smooth control. Despite their complementary roles, integrating asynchronous execution with gradient-based online RL remains underexplored. We present SmoothRL, an online RL framework that fine-tunes a pretrained policy within an asynchronous inference loop. SmoothRL follows a value-gradient paradigm, directly updating policy parameters using gradients of the action-value function with respect to policy actions. To enable correct optimization under asynchronous execution, SmoothRL explicitly models the asynchronous inference process during training. Specifically, each generated action chunk is partitioned by frame index into three regions: a committed region, consisting of actions committed by the previous inference cycle; an execution region, containing newly generated actions executed by the robot; and a discarded region, containing actions superseded by the next inference cycle. Gradients are propagated only through the execution region, ensuring policy optimization aligns with the trajectory distribution induced by asynchronous execution. We evaluate SmoothRL on real-world robotic tasks requiring high precision, as well as highly dynamic tasks that necessitate asynchronous execution.