arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

通过 VLAC-Cut 引导的管道最大化大规模机器人训练后人类效率

HELP: Human-Efficient Large-Scale Robot Post-Training with Rollout Segmentation

Shaopeng Zhai, Qi Zhang, Tianyi Zhang, Haoran Zhang, Fuxian Huang, Zhanhui Lin, Zijun Xu, Weinan Zhang

arXiv 2607.09776首次发表:更新:

发表机构

Shanghai AI Lab(上海人工智能实验室)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究旨在最大化机器人训练后人类效率,提出含专门分工的高效训练后管道,引入 VLAC-CUT 工具提高数据利用效率,经四个实际任务验证,最终策略成功率达 80% - 95%,吞吐量提升 1.7 倍至 4.2 倍,优于仅人工参与训练。

AI 中文摘要

在将视觉语言动作(VLA)模型应用于下游任务时,由于一轮数据无法解决所有问题,需要多轮训练后调整,这使得持续迭代以逐步解决前一轮暴露的弱点成为必要。本报告旨在最大化训练后的人类效率,即每单位人力和时间实现的策略改进和任务吞吐量。我们提出了一种高效的训练后管道,使少数人类操作员能够监督多个机器人。该管道围绕专门的分工构建:训练有素的远程操作员专注于高价值的远程干预和恢复演示,而现场操作员监控多个机器人、触发接管并进行物理重置。这种角色专业化减少了任务切换,降低了操作员培训成本,并允许有限的人力监督更大机群中更多的机器人交互。为了提高数据利用效率,我们引入 VLAC-CUT 作为自动展开筛选工具。它将自主机器人轨迹分为进展、空闲、故障诱导和恢复部分,保留有用部分,同时过滤有害或无信息的部分。经过筛选的展开数据与人工参与数据相结合,用于下一轮训练后调整。我们在四个实际操作任务上验证了所提出的管道。在迭代的训练后调整轮次中,最终策略的成功率达到 80% - 95%,任务吞吐量比基础模型提高了 1.7 倍至 4.2 倍。在相同的人工干预预算下,VLAC-CUT 引导的展开重用在成功率和吞吐量方面均优于仅人工参与训练。

英文摘要

When adapting Vision Language Action (VLA) models to downstream tasks, multiple rounds of post-training are often required to progressively address policy weaknesses. In this report, we focus on maximizing human efficiency during this iterative process, measured by policy improvement and task throughput per unit of human labor and time. We propose HELP, a Human-Efficient Large-scale robot Post-training pipeline in which two specialized operators supervise twelve robots concurrently. A trained Teleoperator provides high-value remote interventions and recovery demonstrations, while a Floor Operator monitors the robot fleet, triggers takeovers, and performs physical resets. This role specialization improves human efficiency by reducing task switching, lowering operator training costs, and expanding robot interaction coverage. Beyond increasing rollout volume, concurrent supervision also broadens the range of policy behaviors observed by the human team, making recurring failure modes easier to identify and enabling more targeted takeovers, resets, and recovery demonstrations. To efficiently utilize the large and mixed-quality rollout data, HELP incorporates \vlac, an automatic rollout segmentation critic specifically designed for this setting. It separates autonomous trajectories into progress-making, idle, failure-inducing, and recovery segments. Useful rollout segments are retained and combined with Human-in-the-Loop data for the next post-training round. Across four real-world manipulation tasks, HELP achieves 80\%--95\% success rates and improves task throughput by 1.7$\times$--4.2$\times$ over the base model. Under matched HITL recovery budgets, VLAC-CUT further amplifies throughput gains by 1.20$\times$--3.43$\times$ and success-rate gains by 1.50$\times$--3.00$\times$ over HITL-only updates.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑