发表机构
The University of Tokyo; RIKEN Center for Advanced Intelligence Project(东京大学; 理化学研究所先进智能项目中心)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出QDOS算法,通过优势加权质量多样性预训练与双数据集复用策略,在分层技能学习框架下,于稀疏奖励任务中显著提升性能,优于强基线算法。
AI 中文摘要
近期研究探索如何利用预先收集的数据集提升强化学习(RL)的策略性能与样本效率,一种有前景的方法是采用两阶段策略:第一阶段从给定数据集中提取多样化技能作为底层策略,第二阶段训练高层策略以解决特定任务。通常,底层策略的提取基于无监督学习,如轨迹变分自编码器(trajectory VAE)。然而,该方法的局限在于底层策略的质量高度依赖数据集的质量。为解决此问题,我们提出QDOS(Quality-Diversity Offline Skill learning,离线技能学习的质量多样性),这是一种用于鲁棒离线到在线学习的统一流程。我们的方法结合了优势加权质量多样性预训练目标,该目标通过每个轨迹片段的估计优势对技能提取和多样性目标进行加权,使模型能够提取多样化且高价值的技能。通过提供鲁棒且与任务相关的技能表示,QDOS显著提升了底层策略所用嵌入技能空间的质量。我们进一步将其与双数据集复用策略相结合,其中离线数据既用于技能预训练,又通过伪标签填充在线经验回放缓冲区。实验表明,QDOS在结构化操作任务和非结构化运动任务中显著优于强基线,证实其在具有挑战性的稀疏奖励领域中加速探索并提升最终回报的能力。
英文摘要
Recent studies investigate how to leverage pre-collected datasets to improve the policy performance and sample efficiency of RL. One promising approach to achieve this goal is to employ a two-stage strategy: In the first stage, diverse skills are extracted as a low-level policy from a given dataset, and a high-level policy is trained to solve a specific task in the second stage. Typically, extraction of the low-level policy is performed based on unsupervised learning such as trajectory VAE. However, a limitation of this approach is that the quality of the low-level policy highly depends on the quality of the dataset. To address this issue, we introduce QDOS (Quality-Diversity Offline Skill learning), a unified pipeline for robust offline-to-online learning. Our approach incorporates an Advantage-Weighted Quality-Diversity pretraining objective, which weights the skill extraction and diversity objectives by the estimated advantage of each trajectory segment. This approach allows the model to extract diverse and high-value skills. By providing robust and task-relevant skill representations, QDOS significantly improves the quality of the embedded skill space used by the low-level policy. We further integrate this with a dual dataset reuse strategy, where offline data is used both for skill pretraining and for populating the online replay buffer via pseudo-labeling. Experiments demonstrate that QDOS significantly outperforms strong baselines in structured manipulation tasks and unstructured locomotion tasks, confirming its ability to accelerate exploration and improve final returns in challenging sparse-reward domains.
CommentsIROS 2026