arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.18108cs.RO

技术报告:用于GR00T N1.7的单步漂移动作头

Technical Report: One-Step Drifting Action Heads for GR00T N1.7

  • ZJU-UIUC Institute, Zhejiang University(浙江大学ZJU-UIUC联合学院)
  • LimX Dynamics(逐际动力)

机构由 AI 辅助整理,请以论文原文为准。

Xihe Shao

AI总结:

本报告研究GR00T N1.7中单步漂移动作头替代迭代扩散头,显著降低推理时间,但导致LIBERO任务成功率系统性下降,呈现速度-成功率权衡。

AI中文摘要:

单步动作生成可以大幅降低视觉-语言-动作(VLA)策略的推理成本,但其对闭环任务成功率的影响仍是一个开放问题。本技术报告研究了一种GR00T N1.7变体,其中迭代扩散-Transformer动作头被替换为单步漂移动作头,并附带一种用于异步块替换的重叠条件扩展。所有多种子漂移运行均在两块NVIDIA A800 GPU上训练。在LIBERO上,该动作头将动作头的平均模型前向时间从约45.3毫秒减少到5.0毫秒,而测得的骨干网络加动作头时间从约70.0毫秒降至30.6毫秒。然而,这种加速伴随着任务成功率的系统性下降。在三个漂移种子下,LIBERO-Spatial上的成功率为64.0±4.0%,LIBERO-Goal上为52.0±1.0%,LIBERO-Long上为26.0±2.6%。较低的种子方差表明,性能下降不能仅由随机初始化解释。我们将该结果报告为速度-成功率权衡,而非整体改进,并讨论了可能的促成因素,包括确定性单步模式平均、批次相关的几何估计、长时间开环块执行,以及同步LIBERO评估未涉及异步重叠路径的事实。

英文摘要:

One-step action generation can substantially reduce the inference cost of vision-language-action (VLA) policies, but its effect on closed-loop task success remains an open question. This technical report studies a GR00T N1.7 variant in which the iterative diffusion-transformer action head is replaced by a one-step drifting action head, together with an overlap-conditioned extension for asynchronous chunk replacement. All multi-seed drifting runs were trained on two NVIDIA A800 GPUs. On LIBERO, the action head reduces the mean model-forward time of the action head from approximately $45.3\,\mathrm{ms}$ to $5.0\,\mathrm{ms}$, while the measured backbone-plus-head time falls from approximately $70.0\,\mathrm{ms}$ to $30.6\,\mathrm{ms}$. However, this speedup is accompanied by a systematic reduction in task success. Across three drifting seeds, success is $64.0\pm4.0\%$ on LIBERO-Spatial, $52.0\pm1.0\%$ on LIBERO-Goal, and $26.0\pm2.6\%$ on LIBERO-Long. The low seed variance indicates that the degradation is not explained by random initialization alone. We report the result as a speed--success trade-off rather than an overall improvement, and discuss likely contributing factors including deterministic one-step mode averaging, batch-dependent geometry estimation, long open-loop chunk execution, and the fact that synchronous LIBERO evaluation does not exercise the asynchronous overlap path.

补充信息

↑