EmbodiedMind:面向高效具身智能的自适应数据整理与前缀树强化学习
EmbodiedMind: Adaptive Data Curation and Prefix-Tree Reinforcement Learning for Efficient Embodied Intelligence
浏览论文内容
中文总结 AI 辅助
针对具身基础模型训练中的样本低效、梯度失衡和信用分配问题,提出数据筛选与分层策略优化范式,含RSFT、IR-GRPO和Trie-GRPO,在18个基准上平均性能达70.02%。
中文摘要 AI 辅助
训练具身基础模型通常需要大规模数据集和大量计算资源,但常常面临三个关键限制:(1)由于低信息量样本导致样本利用效率低下;(2)异构任务间梯度贡献不均衡;(3)长时程规划中严重的信用分配问题,其中轨迹级奖励不加区分地惩罚所有令牌。为解决这些问题,我们提出了一种高效的训练范式,通过策略性数据选择和分层策略优化,实现了最先进的平均性能。我们的方法包含三个协同阶段。首先,基于拒绝采样的微调(RSFT)过滤低信息量样本,以建立稳健的行为先验,同时防止分布崩溃。其次,迭代拒绝GRPO(IR-GRPO)采用按难度分层的任务特定队列,在强化学习迭代中保持数据集平衡,并配合混合奖励机制实现精确的跨任务反馈。第三,为增强长时程任务规划,我们引入了Trie-GRPO,一种基于动作前缀树的新型强化学习算法,能够实现步骤级优势估计。通过将中间正确决策与下游错误隔离,解决了信用分配问题,同时与传统搜索树相比,有效平衡了探索效率和深度。最终,EmbodiedMind在18个基准上实现了70.02%的最先进平均性能,并在长时程任务规划准确性上显著优于其他具身基础模型。我们的项目将发布以供复现。
英文摘要
Training embodied foundation models typically requires massive-scale datasets and extensive computational resources, yet often suffers from three critical limitations: (1) inefficient sample utilization due to low-informative samples; (2) imbalanced gradient contributions across heterogeneous tasks; and (3) severe credit assignment problem in long-horizon planning, where trajectory-level rewards indiscriminately penalize all tokens. To address these issues, we propose an efficient training paradigm that achieves state-of-the-art average performance through strategic data selection and hierarchical policy optimization. Our approach consists of three synergistic stages. First, Rejection Sampling-based Fine-Tuning (RSFT) filters out low-informative samples to establish robust behavioral priors while preventing distributional collapse. Second, Iterative Rejection GRPO (IR-GRPO) employs task-specific queues stratified by difficulty to keep datasets balanced across reinforcement learning iterations, coupled with a hybrid reward mechanism for precise cross-task feedback. Third, to enhance long-horizon task planning, we introduce Trie-GRPO, a novel reinforcement learning algorithm based on action prefix trees, which enables step-level advantage estimation. This resolves the credit assignment problem by isolating intermediate correct decisions from downstream errors, while effectively balancing exploration efficiency and depth compared to conventional search trees. As a result, EmbodiedMind achieves a state-of-the-art average performance of 70.02% across 18 benchmarks, and significantly outperforms other embodied foundation models in long-horizon task planning accuracy. Our project will be released for reproducibility.
发表机构
- ZTE Corporation(中兴通讯股份有限公司)
机构由 AI 辅助整理,请以论文原文为准。