超越噪声引导:面向生成式机器人策略的双潜空间强化学习
Beyond Noise Steering: Dual-Latent Space Reinforcement Learning for Generative Robot Policy
浏览论文内容
中文总结 AI 辅助
本文提出双潜空间强化学习框架,通过预测初始噪声和动作表征两个潜变量,在冻结生成器内部实现表征级控制,无需更新基础策略即可加速在线机器人策略适应。
中文摘要 AI 辅助
预训练的生成式机器人策略从示范中学习富有表现力的动作先验。然而,现有的强化学习方法仅引导噪声空间,却未能在生成过程中调节中间动作表征,导致性能下降和效率低下。为解决这一局限,我们提出了一种新颖的双潜空间强化学习(DLSRL)框架,该框架在冻结生成器内部以表征级控制补充初始噪声引导。具体而言,我们的演员网络预测两个不同的潜变量:一个引导行为生成的初始噪声潜变量,以及一个用于中间特征调制的动作表征潜变量。此外,该表征潜变量被映射为适配器特征,并通过残差连接巧妙注入中间动作令牌的隐藏状态。我们的双控制设计使得无需更新基础策略即可直接调整动作表征。在多种生成式策略架构和机器人操作任务上的实验表明,DLSRL有效加速了在线机器人策略适应,并取得了有竞争力的性能。我们的代码可在 \n\n 此 https URL \n\n 获取。
英文摘要
Pretrained generative robot policies learn expressive action priors from demonstrations. However, existing reinforcement learning methods only steer the noisy space but fail to modulate intermediate action representations during the generation process, resulting in performance degradation and inefficiency. To address this limitation, we propose a novel Dual-Latent Space Reinforcement Learning (DLSRL) framework, which complements initial-noise steering with representation-level control inside the frozen generator. Specifically, our actor network predicts two distinct latent variables: an initial-noise latent variable that steers behavior generation, and an action-representation latent variable for intermediate feature modulation. Moreover, this representation latent variable is mapped to adapter features and ingeniously injected into the hidden states of intermediate action tokens via residual connections. Our dual-control design enables direct adjustment of action representations without updating the base policy. Experiments across generative policy architectures and robotic manipulation tasks show that DLSRL effectively accelerates online robot policy adaptation and achieves competitive performance. Our code is available at \href{https://github.com/xianchaoxiu/DLSRL}{https://github.com/xianchaoxiu/DLSRL}.
发表机构
- Shanghai University(上海大学)
机构由 AI 辅助整理,请以论文原文为准。