arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.19659cs.ROcs.LG

EmbodiedMind:面向高效具身智能的自适应数据整理与前缀树强化学习

EmbodiedMind: Adaptive Data Curation and Prefix-Tree Reinforcement Learning for Efficient Embodied Intelligence

Feifan Wang, Zongbing Zhang, Yu Zhang, Lingfeng Wang, Yurui Zhu, Jin Deng, Mingliang Zhang, Zhengguang Gao, Yongcheng Wang, Jin Xu, Ri Yang

首次发表
浏览论文内容

中文总结 AI 辅助

针对具身基础模型训练中的样本低效、梯度失衡和信用分配问题,提出数据筛选与分层策略优化范式,含RSFT、IR-GRPO和Trie-GRPO,在18个基准上平均性能达70.02%。

中文摘要 AI 辅助

训练具身基础模型通常需要大规模数据集和大量计算资源,但常常面临三个关键限制:(1)由于低信息量样本导致样本利用效率低下;(2)异构任务间梯度贡献不均衡;(3)长时程规划中严重的信用分配问题,其中轨迹级奖励不加区分地惩罚所有令牌。为解决这些问题,我们提出了一种高效的训练范式,通过策略性数据选择和分层策略优化,实现了最先进的平均性能。我们的方法包含三个协同阶段。首先,基于拒绝采样的微调(RSFT)过滤低信息量样本,以建立稳健的行为先验,同时防止分布崩溃。其次,迭代拒绝GRPO(IR-GRPO)采用按难度分层的任务特定队列,在强化学习迭代中保持数据集平衡,并配合混合奖励机制实现精确的跨任务反馈。第三,为增强长时程任务规划,我们引入了Trie-GRPO,一种基于动作前缀树的新型强化学习算法,能够实现步骤级优势估计。通过将中间正确决策与下游错误隔离,解决了信用分配问题,同时与传统搜索树相比,有效平衡了探索效率和深度。最终,EmbodiedMind在18个基准上实现了70.02%的最先进平均性能,并在长时程任务规划准确性上显著优于其他具身基础模型。我们的项目将发布以供复现。

英文摘要

Training embodied foundation models typically requires massive-scale datasets and extensive computational resources, yet often suffers from three critical limitations: (1) inefficient sample utilization due to low-informative samples; (2) imbalanced gradient contributions across heterogeneous tasks; and (3) severe credit assignment problem in long-horizon planning, where trajectory-level rewards indiscriminately penalize all tokens. To address these issues, we propose an efficient training paradigm that achieves state-of-the-art average performance through strategic data selection and hierarchical policy optimization. Our approach consists of three synergistic stages. First, Rejection Sampling-based Fine-Tuning (RSFT) filters out low-informative samples to establish robust behavioral priors while preventing distributional collapse. Second, Iterative Rejection GRPO (IR-GRPO) employs task-specific queues stratified by difficulty to keep datasets balanced across reinforcement learning iterations, coupled with a hybrid reward mechanism for precise cross-task feedback. Third, to enhance long-horizon task planning, we introduce Trie-GRPO, a novel reinforcement learning algorithm based on action prefix trees, which enables step-level advantage estimation. This resolves the credit assignment problem by isolating intermediate correct decisions from downstream errors, while effectively balancing exploration efficiency and depth compared to conventional search trees. As a result, EmbodiedMind achieves a state-of-the-art average performance of 70.02% across 18 benchmarks, and significantly outperforms other embodied foundation models in long-horizon task planning accuracy. Our project will be released for reproducibility.

发表机构

  • ZTE Corporation(中兴通讯股份有限公司)

机构由 AI 辅助整理,请以论文原文为准。

↑